Compare commits

...
Author SHA1 Message Date
ChuckBuildsandClaude Opus 5.5 3ea1fd42df perf(espn): remember settled day chunks between window refreshes
Since ESPN started rejecting date ranges, every scoreboard's hourly
Recent/Upcoming refresh re-asks its 22-day window (14 back, 7 ahead) one
day at a time. Measured on hdpi 2026-10-02 (NFL, college football, MLB,
college baseball, NHL): the hourly refresh was ~270 of 321 ESPN requests
and ~21 of 24.6MB in the hour. Days that ended three or more days ago
cannot change, and they were 68% of the window's bytes (6.9 of 10.2MB).

_fetch_one_chunk now keeps a settled chunk (last day <= UTC today - 3) in
memory for 24h as zlib-compressed JSON, keyed by URL, the other params
and the chunk, and answers it from there. Both range paths go through it:
fetch_espn_scoreboard and BackgroundDataService._fetch_in_date_chunks.
Each hit is parsed afresh, failed and capped chunks are not stored, and
the memory is bounded (512 entries / 8MB compressed). Against live ESPN,
a second refresh of the five hdpi windows went from 111 requests and
10.18MB to 50 requests and 3.23MB with identical events; the memory held
60 entries in 585KB.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:24:59 -04:00
ChuckandClaude Opus 5.5 07abd87d5e fix(sports): scroll and Vegas cards name the printed date's own weekday (#747)
With scroll_card.date_format "weekday", a Friday 8 PM ET game read
"Sat Oct 2" on the scroll and Vegas cards.

Cause: the extractor prints the "M/D" in the plugin's resolved zone (its
own setting, then the global one, then the system zone), but the card is
handed only the plugin's config. Its timezone ships as "", so
card_tzinfo fell back to UTC and the weekday belonged to the UTC date:
the next day for evening games in the Americas, the previous day for
morning games east of UTC (Auckland, Kiritimati).

Fix: every zone is within a day of UTC, so the printed date is the
start's UTC date or a neighbour of it. _format_date_as now takes the game
and names the weekday of whichever of those days has the printed month
and day, falling back to the zone-based weekday only when the start
cannot place the date (no offset, unparseable, or more than a day away).
The switch-mode scorebug shares the formatter and passes the game too, so
the twins stay identical; it already used the resolved zone and draws
what it drew before. Public signatures are unchanged.

Tests cover US DST end, New Year's Eve, both sides of the date line, NZ
DST start and UTC+14. The twins test's weekday pin is updated: the drawn
date now agrees, and only the bare weekday helpers still differ.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:19:05 -04:00
ChuckandClaude Opus 5.5 e32d177cbd fix(web-ui): Plugin Manager - enable aliased installs, Update All, on-demand modes, long installs, categories, GitHub-URL install (#746)
* fix(web-ui): Update All sends the live installed list and redraws the grid

updateAll() preferred PluginStateManager.installedPlugins over
window.installedPlugins. Only updateAll's own end-of-run refresh ever
fills PluginStateManager, so from the second run on it sent the first
run's plugins: one uninstalled since failed with "plugin not found" and
one installed since was never updated. That refresh also only replaced
window.installedPlugins, so the installed cards and the Updates badge
kept offering "Update to vX" for what had just been updated.

Read window.installedPlugins, the list plugins_manager.js republishes
after every install, uninstall and refresh, keeping PluginStateManager
as the fallback for a page without it, and refresh through
pluginManager.loadInstalledPlugins(true), which redraws the grid.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): list each plugin's display modes in /plugins/installed

The on-demand modal fills its Display Mode select from
plugin.display_modes, but /plugins/installed never sent the field. Every
plugin offered one option, its own id, under "This plugin exposes a
single display mode"; the display resolved that id to the plugin's first
mode, so a multi-mode plugin could only be started, or pinned, there.

Add display_modes to each entry, read from the plugin catalog
(get_plugin_display_modes), the same declared list /display/modes and
on-demand/start use, keeping only strings. Single-mode plugins still get
one option and the same hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): enable a store install by its installed id, and not on reinstall

The store's Install button enabled the new plugin by the registry id it
installed. Weather, Music, Stocks and Leaderboard install under the id
their manifests declare (weather -> ledmatrix-weather); the plugin list,
the config section and /plugins/toggle know only that id, so the toggle
answered 404 "Plugin not found" and the plugin stayed disabled behind
"installed, but enabling it failed". The same button on an installed
plugin (Reinstall) enabled it too, switching a plugin the user had
turned off back on.

POST /plugins/install now names the installed plugin: plugin_id in the
direct answer and in the queued operation's result, read from the
installed manifest found the way the store's update and uninstall find
it (_find_plugin_path: id, aliases, plugin_path name), else the
requested id. The client reloads the list, then enables that id; from
an answer without it, the installed entry the store entry matches
(findInstalledStorePlugin, which isStorePluginInstalled now uses). A
reinstall, decided by the same match that labelled the button, reloads
the list and leaves the enabled state alone.

test/js/plugins_manager_sandbox.js runs the whole of
plugins_manager.js in a vm context against a fake DOM and API, for
suites that drive its real flows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): wait for long store installs; on timeout reload, not fail

pollOperationStatus gave a queued install 60 polls, a second apart,
then reported "Install operation timed out" as an error and stopped.
The server allows the plugin's dependency install 300 s on its own
(install_requirements_file in store_install.py), after a download that
fetches the plugin a file at a time, so installs that went on to
succeed were reported as failed, never enabled, and left out of the
installed list until the page was reloaded.

Give installs INSTALL_POLL_MAX_ATTEMPTS (600, ten minutes). When even
that runs out, reload the installed list and the store badges and warn
that the install may still be running; nothing is enabled without the
operation's answer. Uninstall keeps the default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): build the store's category filter from the store's plugins

The #plugin-category select listed seven fixed categories while the
registry uses about twenty (productivity, utility, transit, finance,
...), so roughly a third of the store could not be filtered to, and
"Financial" missed the plugin filed under "finance".

The template now ships only "All Categories"; syncStoreCategoryOptions,
run by applyStoreFiltersAndSort, adds one option per category the cached
store plugins have (case folded, as the filter compares), keeps the
current choice, and rebuilds only when the set changes or the partial
was swapped in afresh -- the way the Starlark section builds its own.

The test sandbox gains window.addEventListener (initPluginsPage needs
it) and quiets the script's "element not found" warnings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): one handler for the GitHub-URL Install button

#install-plugin-from-url had an inline onclick calling
window.handleGitHubPluginInstall, and attachInstallButtonHandler also
gave it a click listener that installs, so both ran on every click
(and on Enter, which clicks it). The inline handler threw a
ReferenceError -- it called isGithubUrl, which is local to the
plugin-manager IIFE, from outside it -- so only the listener's request
went out; correcting that scope alone would have sent every install
twice.

Remove the inline onclick and the window.handleGitHubPluginInstall it
called, which nothing else uses. The listener, which already sent the
only request, is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:54 -04:00
ChuckandClaude Opus 5.5 d18e4d3c9d fix(plugins): sub-package reload, symlinked dev plugins, BaseException in update(), config callbacks outside the lock (#741)
* fix(plugins): drop a plugin's package modules when it unloads

A plugin that keeps helpers in a package (providers/feed.py, imported as
`from providers.feed import ...`) leaves dotted entries in sys.modules.
PluginLoader only tracked bare names: `providers` was namespaced and
dropped on unload, `providers.feed` stayed. A reload after a store update
imported a fresh `providers`, then got the old `feed` back from the module
cache, so the new manager.py ran against the old helpers until the display
restarted. A load that failed part-way left them behind the same way.
Elections (providers/), flights (enrichment/) and olympics (data/,
renderers/) ship packages.

The loader now records the dotted modules whose file (or, for a namespace
package, every __path__ entry) lies inside the plugin directory. They keep
their names while the plugin runs, as before, and unregister_plugin_modules()
drops them, only while sys.modules still holds that plugin's module. The
failed-load cleanup in load_module() drops them too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): remove a symlinked dev plugin as a link

PluginStoreManager._safe_remove_directory, behind uninstall and behind
discarding the set-aside copy after an install or update, handed a
symlinked dev plugin (scripts/dev/dev_plugin_setup.sh) to shutil.rmtree,
which refuses a symlink. The chmod fallback then walked through the link
and set every directory and file in the linked checkout to 0700, and the
sudo stage refused the resolved path as outside the plugins directory. The
removal failed, the link stayed, and the developer's checkout lost its
group/other permissions. A dangling link read as already removed, because
exists() follows it, and was left behind.

A symlink is now unlinked before any other stage runs, and before the
exists() check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): load a dev plugin linked in under a different name

contained_plugin_dir(), the containment check before a plugin's
dependencies are installed, resolved the plugin directory and looked for
the resolved folder's name among the plugins directory's entries. A dev
plugin symlinked in under its id by a name its checkout does not share --
`dev_plugin_setup.sh link-github foo <url>` clones ledmatrix-foo, the
repository naming convention, and links it as plugins/foo -- has no such
entry, so install_dependencies() returned False and the load failed with
"Dependency installation failed", even with no requirements.txt.

When the path sits directly in the plugins directory, the entry it names
(the link) is looked up first; anything else is resolved and matched by
name as before. The answer is still always rebuilt from a name os.scandir()
returned for the plugins directory, so a path outside it is still refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): release a plugin whose update() raises a BaseException

On the async update worker, the wrapped update() finished its bookkeeping
(_finish: release the plugin lock, drop the pending slot, state back to
ENABLED) only for an Exception. asyncio.CancelledError and SystemExit
derive from BaseException, so one raised from update() skipped _finish:
the plugin kept its lock and stayed RUNNING for the life of the process,
never rescheduled, with every display() skipped as busy. PluginExecutor
caught only Exception as well, so its thread died with the call never
marked complete and an immediate failure was logged and recorded as a
timeout.

_target_update now runs _finish for any BaseException and re-raises it,
and the executor's thread stores it like any other exception, so it is
reported as the operation's failure (PluginError) on both the async and
the synchronous path. _finish and _record_update_failure take a
BaseException.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(config): notify config subscribers outside the service lock

ConfigService._load_config ran every subscriber while holding _lock. The
display's per-plugin subscriber calls PluginManager.apply_config_change,
which waits up to PLUGIN_LOCK_TIMEOUT (5 s) for a plugin busy in update().
A save that enables or disables a plugin also flags a reconcile, which the
render thread runs: its get_config(), and the unsubscribe() of a plugin it
disables, both take _lock, so the panel froze behind every slow callback,
up to 5 s per busy plugin.

The config is now swapped under _lock and the subscribers are called after
it is released, from a copy of the subscriber lists. A separate
_notify_lock is held across a whole reload (read, swap, notify), so one
reload's notifications still finish before the next one's start. Each
callback is checked against the live lists just before it runs, and
unsubscribe() waits only for a call of that same callback already in
progress (unless it is that callback's own thread), so a callback it
removed is not running and will not run once it returns, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:42 -04:00
ChuckandClaude Opus 5.5 04f0d8d134 fix(ipc): reset the state subscription's reconnect wait after a good connection (#740)
StateSubscription._run reset its backoff only when _follow() returned
normally, which happens only on stop(). Every real disconnect raises
ControlError, so the wait kept doubling across connections: after
successive display restarts the web resubscribed 1, 2, 4, 8, 16 and then
30 s later for good, answering from one-shot state.get connections in the
meantime. The docs promise "1 s up to 30 s" per outage.

The wait now goes back to the minimum once a connection got as far as
storing a snapshot, whatever ended it. A display without the stream
(unknown_command) is still retried at the slow interval.

The frozen-timestamp bug found in the same review is fixed by #737.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:31 -04:00
ChuckandClaude Opus 5.5 064b9c9912 fix(display): a non-numeric plugin duration no longer stops the display; narrow scroll strips no longer raise (#739)
* fix(display): a plugin duration that is not a number no longer stops the display

DisplayController._get_display_duration returned whatever the plugin's
get_display_duration() gave back. clock-simple, calendar and countdown
return their display_duration setting straight from config.json, so a
value saved as "20" or null reached _resolve_durations as a string or
None, and its `<= 0` check raised a TypeError. Nothing in the loop caught
it: run()'s outer handler logged "Unexpected error in display controller"
and cleanup() ended the service when that plugin's screen came up, and
systemd restarted it into the same crash.

The plugin's answer is now read as seconds: a finite number or a numeric
string is used (as BasePlugin.get_display_duration already accepts), a
number at or below zero still goes to _resolve_durations' 15 s rule, and
anything else -- None, a non-numeric string, a bool, NaN, infinity, or a
get_display_duration() that raises -- gets the 30 s a mode without a
plugin gets. The warning is logged once per plugin, not at every screen.

Tests: test/test_display_duration_not_a_number.py, including the real
run() on the run-loop harness, which returned at t=30 before the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(scroll): a strip narrower than the panel no longer raises on every frame

ScrollHelper._get_visible_portion_integer handled a frame that runs off
the end of the strip by copying the strip's tail and then the rest of the
frame from its head, which assumed the head was at least that wide. For a
strip narrower than the panel that raised "could not broadcast input
array" at every position, so get_visible_portion() never returned a frame
and the caller logged a traceback each frame. Vegas composes such a strip
(lead_in_width defaults to 0) when its content is narrower than the chain.

A wrapping frame is now taken column by column modulo the strip's width
(np.take, mode='wrap', into the reused frame buffer): the tail then the
head, as before, and a narrow strip repeated across the panel. The same
path takes a position before the start of the strip, whose [-n:m] slice
was empty and made frombytes raise; the integer and sub-pixel fast paths
now leave a negative start to it. A zero-width strip is still a black
frame.

Tests: test/test_scroll_helper_narrow_strip.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:19 -04:00
ChuckandClaude Opus 5.5 a6e9e3ef1c fix(cache): cache keys too long to be a filename; memory hits judged by the record's own age (#738)
* fix(cache): store keys too long to be a filename

The calendar plugin's cache key joins every calendar id the user picked.
On hdpi it passed 300 bytes; ext4 caps a filename at 255, so every write
(the temp file, the direct-write fallback and the home-directory fallback)
failed with ENAMETOOLONG, once an hour, and the final warning said
"(permission denied)" whatever the error was.

DiskCache.get_cache_path keeps a key of up to 200 UTF-8 bytes as its
filename, exactly as before, and turns a longer one into its first 183
bytes (cut on a character boundary) plus a 16-hex-digit hash of the whole
key. The temp file adds 15 bytes, so the longest name is 215. The
shortened stem is itself short, so the web UI's cache list, which names a
key by its filename, deletes the same file. The give-up warning now names
the real error.

Validated on ledpi's ext4: the old module drops the hdpi-shaped key, the
new one writes a 205-byte filename and reads it back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cache): judge a memory hit by the record's own timestamp

A record loaded from disk went into the memory tier timed from the load,
so get(key, max_age=300) could return data close to 600 s old: after a
restart, after the memory sweep, or in a second process. A stored ttl was
stretched the same way. #728's _fresh_cached works around it for the
scoreboard; every other caller was exposed.

get_cached_data and load_cache now also check a memory hit against the
record's embedded timestamp, with DiskCache.get's rule that a stored ttl
wins over max_age. A stale copy is dropped and the read falls through to
disk, which returns the other process's newer write if there is one.
Records without a timestamp keep the memory tier's own clock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:00 -04:00
ChuckandClaude Opus 5.5 41192b9588 fix(ipc): ticks carry the volatile timestamps, so current-status stays known (#737)
State stream ticks carry the volatile timestamps (display.last_updated, plugins.published_at), so current-status and the plugin runtime stay fresh while one mode stays on screen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:00:07 -04:00
ChuckandClaude Opus 5.5 ef69201770 perf: import package re-exports on first use (#724)
src/common/__init__.py and src/plugin_system/__init__.py resolve their re-exports lazily (PEP 562 __getattr__, __all__ and __dir__ unchanged, TYPE_CHECKING imports for mypy), and sync_manager imports numpy only where send_frame uses it. The web process no longer loads numpy, freetype helpers and PluginManager just to import path_safety, store_manager or schema_manager (~67 MB to ~54 MB RSS on a Pi 4). from src.common import X and from src.plugin_system import X keep working, including submodule imports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:38:27 -04:00
ChuckandClaude Opus 5.5 ed0753c1ca perf: cheap per-frame and per-fetch savings (#725)
Six small savings with no behaviour change: the odds fetch no longer pretty-prints every response for a debug line; the scroll integer-slice path drops a redundant full-frame np.ascontiguousarray; ledmatrix-web.service gets MALLOC_ARENA_MAX=2 like the display unit; core ESPN responses are parsed via response_json (orjson when installed); and the scroll frame stats go to INFO only for degraded windows plus a 5-minute heartbeat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:26:03 -04:00
ChuckandClaude Opus 5.5 7bb85c0356 perf(cache): skip rewriting unchanged CacheManager.set() records (#730)
The disk cache's unchanged-payload skip now ignores a CacheManager.set() record's timestamp, so unchanged re-saves are skipped; a skip moves the file's mtime to the new timestamp instead, and readers take a record's age from the newer of the two (never more than an hour past the embedded timestamp). Per-plugin plugin_metrics:<id> records become one plugin_metrics_snapshot written at most once a minute, and CacheManager builds its ConfigManager on first use. On hdpi, cache file writes went from ~37 to 8.6 a minute.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:14:08 -04:00
ChuckandClaude Opus 5.5 0d179fdf12 perf(display): throttle the per-frame update tick; check strips without building them (#731)
The frame loops and the dwell sleep run PluginManager.run_scheduled_updates() at most every 0.25 s instead of after every frame (the top of each loop pass still ticks unthrottled), and SportsScrollDisplay and the sync follower ask ScrollHelper.has_strip() instead of building cached_image just to see whether a strip exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:01:23 -04:00
ChuckandClaude Opus 5.5 ee789775b4 perf(install): build rpi-rgb-led-matrix with a faster SetImage (#736)
first_time_install.sh applies patches/rpi-rgb-led-matrix/0001-bulk-setimage.patch just before building the Python binding and reverts it afterwards (and from the EXIT trap), so the submodule stays at its pinned commit. The patch copies each image row with one bulk FrameCanvas::SetPixels call and writes the bit planes branch-free: on a 512x64 Pi 4 the frame copy went from 6.57 ms to 2.21 ms. A patch that no longer applies is reported and skipped. Existing installs get it on a rebuild (RPI_RGB_FORCE_REBUILD=1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:48:27 -04:00
ChuckandClaude Opus 5.5 05deb1ee7d feat(install): support Raspberry Pi OS Bookworm (Python 3.11) alongside Trixie (3.13) (#689)
The installer and scripts/check_system_compatibility.sh share one set of OS rules (scripts/install/lib_os.sh): Bookworm (Debian 12, Python 3.11) and Trixie (Debian 13, Python 3.13) are supported, python3 older than 3.11 stops the install before anything changes, and dhcpcd gets a warning with directions. setcap targets /usr/bin/python3, the apt fallback honours the requirement floors, and the desktop check no longer misreads under pipefail. CI runs the unit and plugin-safety suites on 3.11 and 3.13 (tooling jobs on 3.13); mypy targets 3.11.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:13 -04:00
ChuckandClaude Opus 5.5 f841fa36b6 feat(install): updates refresh systemd units; new installs run the newest release (#729)
Updates that move HEAD now install changed systemd units through a root-owned helper (/usr/local/sbin/ledmatrix-refresh-units, two literal sudo lines), with a backup restored on rollback; a refresh that fails part-way puts the old units back. Devices without the new sudo rule keep updating and are told to re-run the installer once. The one-shot installer now checks out the newest vX.Y.Z release (LEDMATRIX_CHANNEL=beta keeps main) and never moves an existing checkout backwards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:09:53 -04:00
ChuckandClaude Opus 5.5 515248b34e feat(ipc): control socket stage 3 - a state stream replaces polled cache keys (#735)
Adds state.get / state.subscribe to the display's control socket (StateHub in src/ipc/server.py). The web interface holds one subscription per process (web_interface/display_state.py) and reads current-status, on-demand status, plugin runtime and /health's display_loop from it, falling back to the cache keys and heartbeat file. While the socket serves readers, display_current_state and plugin_runtime_snapshot are written less often (about 1.5 instead of 5 cache writes a minute for 15 s screens).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 12:56:47 -04:00
ChuckandClaude Opus 5.5 6bf7c3fa51 perf(layout): id-keyed fit_image cache entries no longer pin the source (#732)
LayoutContext.fit_image keyed images without a cache_key by id() and held
a strong reference to the source so the id could not be recycled. A
plugin following the documented one-liner -- draw_image(Image.open(path),
box) each frame -- never hit that cache and kept the last 64 sources
alive: ~64MB for 500x500 RGBA team logos (median size under
assets/sports), up to ~600MB for the largest.

The entry now holds a weak reference whose callback drops it when the
source is freed, and a hit re-checks that the referent is the same
image. Sources that cannot be weak-referenced are still pinned. Keyed
entries (the only kind any plugin on ledmatrix-plugins main uses today:
football-scoreboard's logo fit) are unchanged.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 12:11:27 -04:00
ChuckandClaude Opus 5.5 76ad71d1c5 fix(frame-timing): keep the GC monitor quiet at interpreter shutdown (#734)
A collection during interpreter shutdown called GcMonitor after the
module's `time` global was torn down, printing "Exception ignored while
calling GC callback ... 'NoneType' object has no attribute
'perf_counter'" at the end of service and test runs.

- GcMonitor binds its clock and sys.is_finalizing at construction and
  does nothing once the interpreter is finalizing.
- install_gc_monitor() unregisters it with atexit; new
  uninstall_gc_monitor().
- DisplayManager.cleanup() (reached from SIGTERM via run()'s finally)
  unregisters it alongside the frame recorder.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 11:51:06 -04:00
124 changed files with 10801 additions and 666 deletions
+3
View File
@@ -10,3 +10,6 @@
# Generated by scripts/build_css.py; collapsed in diffs, not hand-edited. # Generated by scripts/build_css.py; collapsed in diffs, not hand-edited.
web_interface/static/v3/tailwind.css linguist-generated=true web_interface/static/v3/tailwind.css linguist-generated=true
web_interface/static/v3/plugin-frame.css linguist-generated=true web_interface/static/v3/plugin-frame.css linguist-generated=true
# Installed as an executable (its shebang runs it) by install_service.sh.
scripts/install/ledmatrix_refresh_units.py text eol=lf
+1 -1
View File
@@ -31,7 +31,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
# No dependencies: the script reads src/__init__.py and CHANGELOG.md only. # No dependencies: the script reads src/__init__.py and CHANGELOG.md only.
- name: Assert the tag, CHANGELOG and src.__version__ agree - name: Assert the tag, CHANGELOG and src.__version__ agree
+19 -8
View File
@@ -14,8 +14,14 @@ permissions:
jobs: jobs:
plugin-safety: plugin-safety:
name: Plugin safety harness + unit tests name: Plugin safety harness + unit tests (Python ${{ matrix.python-version }})
runs-on: ubuntu-latest runs-on: ubuntu-latest
# The two Pythons the installer supports: Raspberry Pi OS Bookworm ships
# 3.11 and Trixie 3.13.
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.13"]
env: env:
# The bundled fixture plugin gives the harness at least one real plugin # The bundled fixture plugin gives the harness at least one real plugin
# to render, and REQUIRE_PLUGINS turns "discovered zero plugins" into a # to render, and REQUIRE_PLUGINS turns "discovered zero plugins" into a
@@ -29,7 +35,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: ${{ matrix.python-version }}
cache: pip cache: pip
- name: Install dependencies - name: Install dependencies
@@ -43,8 +49,13 @@ jobs:
pytest --no-cov test/plugins/ pytest --no-cov test/plugins/
unit-tests: unit-tests:
name: Core unit tests name: Core unit tests (Python ${{ matrix.python-version }})
runs-on: ubuntu-latest runs-on: ubuntu-latest
# Bookworm's Python (3.11) and Trixie's (3.13); see plugin-safety.
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.13"]
steps: steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2 - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with: with:
@@ -52,7 +63,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: ${{ matrix.python-version }}
cache: pip cache: pip
- name: Install dependencies - name: Install dependencies
@@ -84,7 +95,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
cache: pip cache: pip
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0 - uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
@@ -123,7 +134,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
# Downloads the pinned standalone Tailwind CLI (SHA-256 checked; no # Downloads the pinned standalone Tailwind CLI (SHA-256 checked; no
# Node), rebuilds static/v3/tailwind.css and plugin-frame.css from the # Node), rebuilds static/v3/tailwind.css and plugin-frame.css from the
@@ -142,7 +153,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
cache: pip cache: pip
# The runtime requirements are installed so mypy sees the real types of # The runtime requirements are installed so mypy sees the real types of
@@ -181,7 +192,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
# Stdlib only; exits 0 whatever it finds. # Stdlib only; exits 0 whatever it finds.
- name: Report method-family drift across the nine scoreboards - name: Report method-family drift across the nine scoreboards
+414
View File
@@ -19,6 +19,231 @@ accepts both, but the store flags the old spelling as deprecated
## Unreleased ## Unreleased
### Cheap per-frame and per-fetch savings
- `BaseOddsManager.get_odds()` no longer pretty-prints every odds response
for a debug line: the `json.dumps(..., indent=2)` calls in the fetch path
and `_extract_espn_data` are guarded with `isEnabledFor(DEBUG)`, and the
other debug f-strings there take %-style arguments. Same messages at DEBUG.
- `ScrollHelper`'s integer frame path (`_get_visible_portion_integer`) takes
`tobytes()` straight from the strip's column slice instead of copying it
with `np.ascontiguousarray()` first; the bytes are identical (a test pins
them). At 512x64 on a Pi 4 the bytes step went from ~45 us to ~21 us a
frame.
- `systemd/ledmatrix-web.service` sets `MALLOC_ARENA_MAX=2`, as
`ledmatrix.service` has since #476. Existing installs pick it up when
`scripts/install/install_service.sh` or `install_web_service.sh` is re-run;
until then the startup drift check reports the web unit as changed.
- `APIHelper.get()`/`post()`, `BaseOddsManager.get_odds()`, the two
`LogoDownloader` team fetches and `DynamicTeamResolver`'s rankings fetch
parse with `src.common.json_body.response_json` (orjson when installed),
like `background_data_service` already did. A body orjson rejects falls
back to `response.json()`, so a bad body raises the same
`requests.exceptions.JSONDecodeError` these call sites already catch.
- The `Scroll frame stats` line is logged at INFO only for a degraded window
(fps under 0.9 of the rate the window was locked to, or more than 1% of
frames stalled), the window after one, and a 5-minute heartbeat per
scroller, as the `Vegas FPS` line already was; every window is still logged
at DEBUG. `docs/SCROLL_PERFORMANCE.md` says how to see them all.
### Fewer SD-card writes from the cache
- **An unchanged `CacheManager.set()` no longer rewrites the file.**
`DiskCache` already skipped a payload identical to the last one it wrote,
but `set()` stamps every record with the current time, so for `set()` the
payload never matched and every unchanged re-save was a full rewrite. The
comparison now leaves out a header-first record's timestamp (the `ttl` and
the data still count), and the newer timestamp is kept in the file's mtime
instead: a skipped save touches the file to the record's timestamp, and a
real write pins mtime to the record's own timestamp. Every reader ages a
record from the newer of the two -- `DiskCache.get`, its header-only
staleness check, and the record it returns, whose `timestamp` is the newer
value, so `CacheManager.get`, the memory tier and plugins reading
`record['timestamp']` all agree; the retention sweep and the web UI's cache
list already used mtime. The mtime is trusted at most an hour past the
record's own timestamp, and unchanged data is rewritten once an hour, so a
file copied without its mtime reads at most an hour fresher than its
contents. 100 identical `set()` calls of a 32 KB record: 100 writes before,
1 after.
- **Plugin metrics are one record, written at most once a minute.** The
resource monitor wrote a `plugin_metrics:<id>` record per plugin, each at
most every 30 s: two writes a minute per plugin, 28 on a fourteen-plugin
rig. Every plugin's metrics now go in one `plugin_metrics_snapshot` record
(`{"schema": 1, "plugins": {id: record}}`, each record shaped as before),
written at most once a minute. `GET /api/v3/plugins/metrics` and
`/plugins/metrics/<id>` return the same fields; the numbers can be up to a
minute old instead of 30 s. A plugin the snapshot does not have yet is
still read from its old `plugin_metrics:<id>` record, which nothing writes
any more and the cache's retention removes. Each write starts from the
snapshot on disk, so plugins the display has not run since a restart keep
their numbers, and a reset from the web UI sticks for a plugin the display
is not running, as it did. A plugin with no call for 30 days is dropped from
the snapshot, as its record used to age out.
- **`CacheManager` no longer loads the config when it is built.** Every
manager built a `ConfigManager` and loaded the whole config for a cache
strategy that stopped reading it. `cache_manager.config_manager` is still
there -- the sports plugins resolve the global timezone through it -- and
is now built and loaded on first access; assigning it still replaces it.
`CacheStrategy` is given no config manager (it reads none).
### Plugin update tick: a few times a second, not every frame
- The frame loops and the dwell sleep ran
`PluginManager.run_scheduled_updates()` after every frame, about 125 times
a second on a scroller. Each pass copies the plugin dict and takes several
locks per plugin, almost always to find nothing due: about 100 us with 20
plugins on a Pi 4, 1.2% of the render thread. They now call
`DisplayController._tick_plugin_updates_if_due()`, which runs the pass at
most every `PLUGIN_UPDATE_TICK_INTERVAL` (0.25 s), so a 4 s scroll runs 16
passes instead of 500. No update interval is shorter than 5 s
(`MIN_DYNAMIC_UPDATE_INTERVAL`), and the 1 Hz frame loop already ticked
once a second, so an update starts at most a quarter second later.
- The top of each loop pass still runs it unthrottled, so a plugin just
loaded, reloaded or enabled for on-demand is updated at once. Vegas's own
update thread (`_tick_plugin_updates_for_vegas`) is unchanged.
`test/test_plugin_update_tick_throttle.py` covers both, on the real
`run()` through the golden-trace harness.
### Strip checks no longer build the PIL image
- `SportsScrollDisplay.display_scroll_frame` (every frame) and
`has_cached_content`, and the sync follower's per-frame check of the Vegas
strip, asked whether there was a strip by reading
`ScrollHelper.cached_image`. After the helper deferred the image (an
append, trim or patch), that read built it from `cached_array` and kept
it: 3.5 ms and about 1 MiB more held for a 4288x64 strip, on top of the
array's 0.8 MiB. They now ask `has_strip()`, which gives the same answer
from the helper's bookkeeping. A scoreboard whose `scroll_helper` has no
`has_strip` (its own helper, a test double) is still asked
`cached_image`.
- A strip built with `create_scrolling_image` or `set_scrolling_image`
still keeps both the image and the array, as before.
### Faster frame copy into the panel (library patch, applied at build time)
- Copying each frame into the panel buffer (`SetImage`) was the biggest CPU
cost LEDMatrix owns on large panels: 6-7.5 ms per frame on a 512x64 Pi 4 at
~85 fps, about 60% of a core. The library's binding walked the image column
by column and set one pixel at a time, and each pixel rewrote a word in
every PWM bit plane, 2KB apart, so nearly every write missed the cache.
`patches/rpi-rgb-led-matrix/0001-bulk-setimage.patch` copies row by row in
one bulk call per row, with the colour lookup done once and branch-free
bit-plane writes. The panel buffer is byte-identical to before (882 checks
across image types, offsets, PWM bits, brightness, inverse colours and a
pixel mapper).
- Measured on hdpi (Pi 4, 4x128x64): frame copy 6.57 -> 2.21 ms, the display
process 139% -> 103% of a core, late frames 7.8 -> 5.4 per 1,000.
- `first_time_install.sh` applies the patch to `rpi-rgb-led-matrix-master`
just before building the binding and takes it back out straight after (and
on any exit), so the submodule stays at its pinned commit with no local
changes. A patch that no longer applies after a submodule bump is reported
and skipped; the unpatched library still builds.
- Existing installs keep the library they have until it is rebuilt:
`sudo RPI_RGB_FORCE_REBUILD=1 ./first_time_install.sh`.
`scripts/build_rgbmatrix_nogil.sh` builds from an unpatched copy and is
unchanged.
### Install
- Raspberry Pi OS **Bookworm** (Debian 12, Python 3.11) is supported,
alongside **Trixie** (Debian 13, Python 3.13). The installer used to stop
on anything but Trixie. Which releases and Pythons are accepted now lives
in one place, `scripts/install/lib_os.sh`, which `first_time_install.sh`
and `scripts/check_system_compatibility.sh` both read, so the two can no
longer disagree (the compatibility check called Bookworm an error, and
still accepted Python 3.10, which the rgbmatrix bindings refuse). An
unsupported system gets plain directions to the right image; a `python3`
older than 3.11 stops the install before anything changes.
- The installer says up front when the Pi runs dhcpcd instead of
NetworkManager, and how to switch back: the web page's WiFi tab and the
`LEDMatrix-Setup` hotspot need NetworkManager. Not fatal, and it does not
switch the network stack itself, since that can cut the SSH session.
- The desktop check no longer misses a desktop install: `dpkg -l | grep -q`
under `pipefail` read a match as "not found".
- `cap_sys_nice` is set on the interpreter the services run
(`/usr/bin/python3`); it preferred `/usr/bin/python3.13` whenever it
existed.
- The Step 7 dependency fallback (`scripts/install_dependencies_apt.py`) no
longer accepts apt packages older than the pins -- Bookworm's Flask 2.2.2
and Pillow 9.4, Trixie's Flask 3.1.1 and Pillow 11.1. The floors are read
from `web_interface/requirements.txt`, and pip is asked for `Pillow`, not
`PIL`.
- CI runs the unit and plugin-safety suites on Python 3.11 and 3.13 (was
3.12); mypy targets 3.11.
### Updates refresh the systemd units; new installs run the newest release
- **Updates now install changed systemd units.** An update (Update Code, or
the weekly automatic update) moved the checkout's `systemd/*.service`
templates but never the units systemd runs, so settings added after a
device was installed -- #687's render-loop watchdog, for one -- only ever
arrived with a reinstall. After an update that moves HEAD, the web
interface compares the installed `ledmatrix.service`,
`ledmatrix-web.service` and `ledmatrix-update-verify.{service,path}` with
the new templates (rendered exactly as `install_service.sh` does, comments
ignored as the startup drift warning does) and, when they differ, runs the
new root-owned helper `/usr/local/sbin/ledmatrix-refresh-units`
(`scripts/install/ledmatrix_refresh_units.py`) through sudo: it installs
the changed units and runs `systemctl daemon-reload`, so the restart that
follows the update runs under them. Update Code's message says so.
- **Rollback restores them.** The helper keeps the units it replaced
(`/var/lib/ledmatrix/unit-backup`, root only); when the automatic update's
health check rolls an update back, it runs `ledmatrix-refresh-units
--restore` before restarting the services onto the old code.
- **The sudo rule needs a reinstall.** `install_service.sh` installs the
helper and `lib_sudoers.sh` grants it with exactly two command lines (no
arguments, and `--restore`). A device installed before this has neither;
its updates keep working, log that the new unit settings need a reinstall
and say so in Update Code's message, the same remedy as the startup
"unit drift" warning. Re-run `sudo ./first_time_install.sh` once (or
`sudo ./scripts/install/install_service.sh` then
`./scripts/install/configure_web_sudo.sh`).
- `install_service.sh` now leaves the units it installs mode `0644`, as
`first_time_install.sh` already did; run on its own it left them `0600`.
- **New installs run the newest release.** The one-shot installer cloned
`main`'s tip, so a new device ran unreleased code until the next release.
It now checks out the newest `vX.Y.Z` tag after cloning (the same semver
rules as `web_interface/update_channel.py`), and that release's own
`first_time_install.sh` runs. `LEDMATRIX_CHANNEL=beta` installs `main`
instead and records the beta channel; `first_time_install.sh --beta` (or
`LEDMATRIX_CHANNEL=beta|stable`) records a channel for a manual install.
- **Re-running the one-shot never moves backwards.** On an existing stable
checkout it moves to the newest release only when that release contains
the current commit; a checkout newer than every release keeps its
fast-forward pull (on a branch) or stays put (detached), and beta keeps the
pull it always had. It used to fast-forward a detached release checkout to
`main`'s tip.
### Control socket stage 3: the display's state over the socket
- Two new commands, still protocol version 1. `state.get` returns a
versioned snapshot of what the display is doing: the current mode and
plugin, the on-demand session, the brightness, the plugin runtime
snapshot and the render loop's heartbeat age. With `since`/`epoch` it
returns a short "unchanged" answer. `state.subscribe` returns the same
snapshot, then pushes a `state` event on every change (always the latest
version) and a `tick` at least every 5 s. The display serves all of it
from memory (`StateHub` in `src/ipc/server.py`), and publishing never
waits for a reader. Subscribers have their own bound (4), separate from
the 8 request slots, and one that stops reading is dropped after the 2 s
IO timeout. See `docs/IPC_CONTROL_SOCKET.md`, "The state stream".
- The web interface holds one subscription per process
(`web_interface/display_state.py`). `/display/current-status`,
`/display/on-demand/status`, the plugin runtime fields of
`/plugins/installed` and `/plugins/state`, the reconciliations and
`/health`'s `display_loop` read it first. When the socket is missing (a
stopped or older display, Windows), they fall back to the cache keys and
the heartbeat file. Each answer has a `source` (`socket`, `cache` or
`heartbeat_file`). The stale and stalled rules from #726 apply the same
way to both.
- Fewer SD-card writes while the socket serves those readers.
`display_current_state` is written once a minute and on a flag change,
not on every mode change. The `plugin_runtime_snapshot` refresh goes from
60 s to 120 s. For a rotation of 15 s screens, that is 1.5 cache writes a
minute instead of 5. Both keys keep being written for one release.
- `RenderWatchdog.liveness()` reports the heartbeat age from memory.
`PluginRuntimeView` has a `source`, and `describe()` includes it.
### Web UI: four more tabs are ES-module pages (stage 2) ### Web UI: four more tabs are ES-module pages (stage 2)
- Rotation, Operation History, Config Editor and Backup & Restore follow the - Rotation, Operation History, Config Editor and Backup & Restore follow the
@@ -56,6 +281,27 @@ accepts both, but the store flags the old spelling as deprecated
`scripts/render_bench.py` records the same. Diagnostic only: nothing tunes, `scripts/render_bench.py` records the same. Diagnostic only: nothing tunes,
freezes or disables the collector. freezes or disables the collector.
### Web interface: lighter package imports
- `src.common` and `src.plugin_system` now import their re-exported names on
first use (PEP 562 module `__getattr__`) instead of in `__init__.py`.
`from src.common import ScrollHelper`, `src.plugin_system.PluginManager`,
`from src.common import *` and every submodule import work as before and
return the same objects. What changes is that importing a submodule --
the web interface's `src.common.path_safety`, `src.plugin_system.store_manager`
and the like -- no longer loads `ScrollHelper`, `LogoHelper`, `APIHelper`,
the adaptive layout helpers and `PluginManager` with it. `sync_manager`
imports numpy inside `send_frame`, the one place it uses it, since the API
blueprint imports that module only for its constants.
- The web process no longer loads numpy at all. On a Pi 4 (Python 3.13),
importing `web_interface.app` went from ~67 MB to ~54 MB RSS and from
~2.8 s to ~1.3 s (`-X importtime`, median of five). A bare
`import src.common` went from ~50 MB / ~0.85 s to ~10 MB / ~30 ms. The
display process loads the same modules as before, only later.
- A misspelt name in `from src.common import ...` still raises `ImportError`.
`test/test_lazy_package_imports.py` checks that the packages import nothing
heavy and that every name in `__all__` resolves to its home module's object.
### Outlined text: one rasterization ### Outlined text: one rasterization
- New `draw_text_outlined(draw, xy, text, font, fill, outline_color=(0, 0, - New `draw_text_outlined(draw, xy, text, font, fill, outline_color=(0, 0,
@@ -242,6 +488,43 @@ policies are unchanged.
### Fixes ### Fixes
- A cache key too long to be a filename is now cached. The calendar
plugin's key joins every calendar id the user picked; on a real install
it passed 300 bytes, ext4 refuses names over 255, and every write failed
with `File name too long` — logged as "(permission denied)", so it read
like a cache-directory ownership problem. `DiskCache.get_cache_path` now
keeps a key of up to 200 UTF-8 bytes as its filename, as before, and
turns a longer one into its first bytes plus a hash of the whole key. The
web UI's cache list and delete keep working, because the shortened name
maps back to the same file. A failed write now names the real error.
- The cache's memory tier no longer serves data older than the reader asked
for. A record loaded from disk was timed in memory from the load, not
from when it was written, so `get(key, max_age=300)` could return data
close to 600 s old (after a restart, after the hourly memory sweep, or in
the other process, which only ever loads the record from disk), and a
stored `ttl` was stretched the same way. A memory hit is now also checked against
the record's own timestamp, and a stale one falls through to disk, which
returns a newer write if there is one.
- The garbage-collection timer (`GcMonitor`, above) no longer prints
`Exception ignored while calling GC callback ... 'NoneType' object has no
attribute 'perf_counter'` when the display service or a test run exits.
A collection during interpreter shutdown called it after the module's
`time` global was torn down. The monitor now binds its clock at
construction and does nothing once `sys.is_finalizing()`;
`install_gc_monitor()` unregisters it with `atexit`, and
`DisplayManager.cleanup()` (reached from SIGTERM through `run()`'s
`finally`) unregisters it with the frame recorder. New
`frame_timing.uninstall_gc_monitor()`.
- The web interface's state subscription (`StateSubscription`,
`src/ipc/client.py`) resubscribes about 1 s after a display restart, every
time. Its reconnect wait went back to the minimum only when the
subscription was stopped. A disconnect after a working connection kept
doubling the wait, so successive display restarts were followed by waits
of 1, 2, 4, 8, 16 and then 30 s for good.
During each wait the web answered from one-shot `state.get` connections
instead of its copy. The wait now resets once a connection has stored a
snapshot. A display that does not offer the stream is still retried
slowly.
- A plugin reload after a store update (`plugin.reload`, #720) no longer - A plugin reload after a store update (`plugin.reload`, #720) no longer
freezes the panel during Vegas. On ledpi a football reload froze it for freezes the panel during Vegas. On ledpi a football reload froze it for
3.0 s (`Render stall over: no frame for 3043ms`). The reload ran on the 3.0 s (`Render stall over: no frame for 3043ms`). The reload ran on the
@@ -258,6 +541,25 @@ policies are unchanged.
for the plugin is refused (`plugin-reloading`), and a config reconcile for the plugin is refused (`plugin-reloading`), and a config reconcile
neither loads it twice nor unloads it mid-load. A Vegas fetch that waited neither loads it twice nor unloads it mid-load. A Vegas fetch that waited
out a reload for the lock skips the old instance. out a reload for the lock skips the old instance.
- A plugin display duration that is not a number no longer stops the
display. Several plugins (clock-simple, calendar, countdown) return their
`display_duration` setting as it is in config.json, so a value saved as
`"20"` or `null` (the raw config editor, a hand edit) reached the run loop
as a string or None. Comparing it with 0 raised a TypeError that no
handler in the loop caught: the display service exited when that plugin's
screen came up, and systemd restarted it into the same crash. The
controller now reads the plugin's answer as a number: a numeric string
counts, and anything else (or a `get_display_duration()` that raises)
shows the mode for 30 s, with one warning per plugin.
- A scroll strip narrower than the panel scrolls instead of raising on every
frame. When a frame ran off the end of the strip, `ScrollHelper` copied
the strip's tail and then the rest of the frame from its head, which
assumed the head was that wide; for a narrower strip that raised
`ValueError: could not broadcast` at every position, so nothing was drawn
and each frame logged a traceback. Vegas builds such a strip, with no
lead-in, when its content is narrower than the chain. A frame that runs
off the strip now continues from its head column by column, so a narrow
strip repeats across the panel; a wide strip wraps exactly as before.
- The schedule-off blank and the WiFi notice no longer start with a - The schedule-off blank and the WiFi notice no longer start with a
scroller's leftovers. Both are drawn by the display controller rather than scroller's leftovers. Both are drawn by the display controller rather than
dispatched to a plugin, so #716's handover never reached them: drawn while dispatched to a plugin, so #716's handover never reached them: drawn while
@@ -267,6 +569,46 @@ policies are unchanged.
scroller or Vegas) were counted as 0.5-1 s freezes and logged as a scroller or Vegas) were counted as 0.5-1 s freezes and logged as a
`Render stall ... mid-scroll`. The controller now ends the scroll state `Render stall ... mid-scroll`. The controller now ends the scroll state
before drawing either. before drawing either.
- A plugin that keeps helpers in a package (elections' `providers/`,
flights' `enrichment/`, olympics' `data/` and `renderers/`) now runs its
updated helpers after a reload. Unloading dropped the package itself but
left its modules (`providers.feed`) in `sys.modules`, so the reload after a
store update imported the new `manager.py` and got the old helpers back from
the cache until the display restarted. `PluginLoader` now drops a plugin's
package modules when it unloads, and when a load fails part-way.
- Uninstalling a dev plugin that `scripts/dev/dev_plugin_setup.sh` linked
into the plugins directory now removes the link and leaves the checkout
alone. The store's removal passed the link to `shutil.rmtree`, which
refuses a symlink; its fallback then walked through the link and chmodded
every directory and file of the linked checkout to 0700, and the sudo stage
refused a path outside the plugins directory, so the uninstall failed with
the link still in place. The same removal discards the set-aside copy after
an install or update. A symlink, dangling or not, is now unlinked.
- A dev plugin linked in under a name its checkout does not share now loads.
`dev_plugin_setup.sh link-github foo <url>` clones `ledmatrix-foo` (the
repository naming convention) and links it as `plugins/foo`. The loader's
containment check for dependency installs resolved the link and looked for
`ledmatrix-foo` among the plugins directory's entries, found none, and
refused the plugin, so the load failed with "Dependency installation
failed" even when it had no `requirements.txt`. The check now looks for the
entry the path itself names in the plugins directory, the link, and still
only ever answers with an entry it found there.
- A plugin whose `update()` raises `asyncio.CancelledError` or `SystemExit`
no longer goes dark until a restart. Both derive from `BaseException`, not
`Exception`, and the update worker's bookkeeping caught only `Exception`:
the plugin kept its lock and stayed RUNNING, so it was never updated again
and every `display()` was skipped as busy. It is now recorded as that
update's failure, the same as any other raise. The plugin executor
reported such a call as a timeout; it now reports it as a failure.
- Saving a config change no longer freezes the panel while a plugin is busy.
`ConfigService` told its subscribers about a change while holding its lock,
and the display's per-plugin subscriber waits up to 5 s for a plugin in the
middle of an update. A save that enables or disables a plugin also queues a
reconcile, which the render thread runs, and its `get_config()` and
`unsubscribe()` waited behind every one of those callbacks. Subscribers now
run after the lock is released. One reload's notifications still finish
before the next one's start, and a callback `unsubscribe()` removed is not
running, and will not run, once it returns.
- A plugin whose `display()` raises now opens its circuit breaker. The first - A plugin whose `display()` raises now opens its circuit breaker. The first
frame of each screen goes through the plugin executor, which caught the frame of each screen goes through the plugin executor, which caught the
exception and returned False. The display read that as "no content" and exception and returned False. The display read that as "no content" and
@@ -304,6 +646,48 @@ policies are unchanged.
being stopped, blanks the panel within about a second. It used to stay on being stopped, blanks the panel within about a second. It used to stay on
until the next minute, because the once-a-minute schedule check had until the next minute, because the once-a-minute schedule check had
already run that minute and the session had overridden its answer. already run that minute and the session had overridden its answer.
- Check & Update All updates what is installed now. A second run in the
same page sent the plugins the first run had seen, so a plugin uninstalled
since then failed with "plugin not found" and one installed since was
skipped. After a run the installed cards and the Updates badge show the
new versions; they kept offering "Update to vX" for what had just been
updated until the page was reloaded.
- The Run On-Demand dialog lists a plugin's display modes, so a mode other
than the first can be started, and pinned. `/api/v3/plugins/installed`
never sent `display_modes`, which the dialog reads, so every plugin
offered only its own id under "This plugin exposes a single display
mode", and the display started its first mode. Each entry now carries
`display_modes`, the modes its manifest declares.
- Installing Weather, Music, Stocks or Leaderboard from the Plugin Store
enables it, as installing any other plugin does. Each installs under the
id its manifest declares (`ledmatrix-weather` for the store's `weather`),
but the store enabled the store id, which `/api/v3/plugins/toggle`
answered with "Plugin not found": the plugin stayed disabled behind
"installed, but enabling it failed". `POST /api/v3/plugins/install` now
answers with the installed `plugin_id` (in the operation's result when it
is queued), and the store enables that.
- Reinstalling a plugin from the Plugin Store leaves it enabled or disabled
as it was. Reinstall enabled it as a fresh install does, so a plugin the
user had switched off came back on.
- A Plugin Store install that takes more than a minute is no longer
reported as failed. The store stopped waiting after 60 s and showed
"Install operation timed out" while the server, which allows the
plugin's dependency install 300 s on its own, carried on and usually
succeeded; the plugin was then neither enabled nor listed until the page
was reloaded. The store now waits up to 10 minutes, and if it still has
no answer it reloads the installed list and says the install may still
be running.
- The Plugin Store's category filter lists every category its plugins
have. It offered a fixed seven while the registry uses about twenty, so
plugins filed under productivity, utility, transit and the rest could not
be filtered to, and "Financial" missed the plugin filed under "finance".
The choices are now built from the store's plugins, as the Starlark
section's are.
- The Install button under Install Single Plugin (Plugin Manager > Install
from GitHub) runs one handler per click. It also had an inline `onclick`
whose handler threw a `ReferenceError` on every click; only the other
handler's request went out, and making the inline one work would have
sent every install twice. The inline handler is gone.
- `/api/v3/plugins/installed` no longer reports the display's plugins as - `/api/v3/plugins/installed` no longer reports the display's plugins as
`live` while `/api/v3/health` says `display_loop: stalled`. The runtime `live` while `/api/v3/health` says `display_loop: stalled`. The runtime
snapshot is written from its own thread, which kept going while the render snapshot is written from its own thread, which kept going while the render
@@ -315,11 +699,41 @@ policies are unchanged.
heartbeat when the service stops), is `stale` at once instead of `live` heartbeat when the service stops), is `stale` at once instead of `live`
for up to 180 s. No new files or writes: both checks are on the reading for up to 180 s. No new files or writes: both checks are on the reading
side. side.
- A scoreboard's scroll and Vegas cards with `scroll_card.date_format:
"weekday"` now show the printed date's own weekday. A Friday 8 PM ET game
read "Sat Oct 2". The card took the weekday in the plugin's own
`timezone` setting, which ships blank, so it fell back to UTC, while the
"Oct 2" beside it came from the zone the plugin actually resolves (its
setting, then the global one, then the system zone). Every zone is within a
day of UTC, so the card now finds which day near the start's UTC date has
the printed month and day and names that one. Games east of UTC (Auckland,
Kiritimati) were off by a day the other way and are fixed the same way.
The switch-mode scorebug, which already used the plugin's resolved zone,
shares the same formatter and draws what it drew before.
- `/api/v3/display/current-status` reflects a wake from scheduled-off, a - `/api/v3/display/current-status` reflects a wake from scheduled-off, a
schedule-off blank, or an on-demand session starting or ending at once, schedule-off blank, or an on-demand session starting or ending at once,
even when the mode name stays the same. The display republished its even when the mode name stays the same. The display republished its
current state only on a mode change or every 30 s, so `is_display_active` current state only on a mode change or every 30 s, so `is_display_active`
and `on_demand_active` could be up to 30 s out of date. and `on_demand_active` could be up to 30 s out of date.
- `/api/v3/display/current-status` no longer answers `mode: null` over the
control socket (#735) once the same mode has been on screen for more than
two minutes: a live game under live priority, Vegas, or a single plugin.
On ledpi it returned nulls in every sample for 90 minutes while the
display was live. The state stream's version leaves out the timestamps
that move on every publish, and a subscriber's keepalive tick carried only
the loop's heartbeat. So the web interface's copy kept the
`display.last_updated` of the last real change, and the reader's 120 s
rule called it unknown. The plugin runtime section had the same problem:
with no plugin changing state, `/plugins/state` and the `runtime` in
`/plugins/installed` read `stale` after 180 s. A tick (and a `state.get`
answer with `since`) now carries `volatile`: the current values of those
timestamps (`display.last_updated`, `on_demand.last_updated` and
`remaining`, `plugins.published_at`), and the subscription merges them
into its copy. The verdicts are unchanged. A render thread that stops
publishing still reads as `stalled` after 60 s and as unknown after 120 s,
a runtime publisher that stops still goes `stale`, and a subscription that
goes quiet still falls back to the cache. The cache path's 120 s rule is
unchanged.
### Scrolling ### Scrolling
+6 -1
View File
@@ -151,6 +151,11 @@ The system supports live, recent, and upcoming game information for multiple spo
- **1GB models (Pi 3B / 3B+), the 512MB Pi Zero 2 W and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`. Once running, keep an eye on memory: see [docs/LOW_MEMORY_BOARDS.md](docs/LOW_MEMORY_BOARDS.md). - **1GB models (Pi 3B / 3B+), the 512MB Pi Zero 2 W and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`. Once running, keep an eye on memory: see [docs/LOW_MEMORY_BOARDS.md](docs/LOW_MEMORY_BOARDS.md).
### Operating system
- **Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12)**, 64-bit recommended. Trixie is the current release and the one to pick for a new SD card; an existing Bookworm install works as it is, no upgrade needed. The installer checks this first and stops with directions on anything else (Bullseye and older, the desktop edition, other distributions).
- **Python**: whatever the OS ships, 3.13 on Trixie and 3.11 on Bookworm. Don't install a different Python; the installer and the services use the system `python3`.
- **Networking**: NetworkManager, the default on both. Choosing a WiFi network from the web page and the `LEDMatrix-Setup` hotspot need it; if you switched to dhcpcd in `raspi-config`, switch back (Advanced Options → Network Config → NetworkManager).
### RGB Matrix Bonnet / HAT ### RGB Matrix Bonnet / HAT
- [Adafruit RGB Matrix Bonnet/HAT](https://www.adafruit.com/product/3211) – supports one “chain” of horizontally connected displays - [Adafruit RGB Matrix Bonnet/HAT](https://www.adafruit.com/product/3211) – supports one “chain” of horizontally connected displays
- [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular` as hardware mapping)* - [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular` as hardware mapping)*
@@ -249,7 +254,7 @@ These are not required and you can probably rig up something basic with stuff yo
<img width="512" height="361" alt="Step 2 Other " src="https://github.com/user-attachments/assets/166a22e8-8067-48df-9f80-50c91f573356" /> <img width="512" height="361" alt="Step 2 Other " src="https://github.com/user-attachments/assets/166a22e8-8067-48df-9f80-50c91f573356" />
5. Then choose Raspbian OS (64-bit) Lite (Trixie) 5. Then choose Raspbian OS (64-bit) Lite (Trixie). Bookworm Lite (listed as Legacy) also works; see [Operating system](#operating-system) below
<img width="512" height="361" alt="Step 4 Trixie Lite 64" src="https://github.com/user-attachments/assets/3b8590ce-b810-4dfe-9253-26e0d4f8ed1e" /> <img width="512" height="361" alt="Step 4 Trixie Lite 64" src="https://github.com/user-attachments/assets/3b8590ce-b810-4dfe-9253-26e0d4f8ed1e" />
+2 -2
View File
@@ -17,13 +17,13 @@ The LEDMatrix emulator allows you to run and test LEDMatrix displays on your com
## Prerequisites ## Prerequisites
### System Requirements ### System Requirements
- Python 3.10 or higher - Python 3.11 or higher (3.11 and 3.13 are tested)
- Windows, macOS, or Linux - Windows, macOS, or Linux
- At least 2GB RAM (4GB recommended) - At least 2GB RAM (4GB recommended)
- Internet connection for plugin downloads - Internet connection for plugin downloads
### Required Software ### Required Software
- Python 3.10+ - Python 3.11+
- pip (Python package manager) - pip (Python package manager)
- Git (for plugin management) - Git (for plugin management)
+19 -1
View File
@@ -15,6 +15,12 @@ This guide will help you set up your LEDMatrix display for the first time and ge
- Power supply (5V, 4A minimum recommended) - Power supply (5V, 4A minimum recommended)
- MicroSD card (16GB minimum) - MicroSD card (16GB minimum)
**Software:**
- Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12). Trixie
is the current release; Bookworm is listed as Legacy in Raspberry Pi
Imager. No other system is supported, and the installer says so up front.
- The OS's own Python: 3.13 on Trixie, 3.11 on Bookworm
**Network:** **Network:**
- WiFi network (or Ethernet cable) - WiFi network (or Ethernet cable)
- Computer with web browser on same network - Computer with web browser on same network
@@ -28,7 +34,8 @@ This guide will help you set up your LEDMatrix display for the first time and ge
There is no prebuilt SD card image — you install LEDMatrix onto stock There is no prebuilt SD card image — you install LEDMatrix onto stock
Raspberry Pi OS Lite yourself: Raspberry Pi OS Lite yourself:
1. Flash Raspberry Pi OS Lite to the MicroSD card (Raspberry Pi Imager) 1. Flash Raspberry Pi OS Lite (Trixie, or Bookworm) to the MicroSD card
(Raspberry Pi Imager)
2. Connect the LED matrix to your Raspberry Pi, insert the card, and 2. Connect the LED matrix to your Raspberry Pi, insert the card, and
power on power on
3. SSH into the Pi and run the one-shot installer: 3. SSH into the Pi and run the one-shot installer:
@@ -39,6 +46,17 @@ Raspberry Pi OS Lite yourself:
[README Installation Steps / Quick Install](../README.md#installation-steps) [README Installation Steps / Quick Install](../README.md#installation-steps)
for full details for full details
The one-shot installer installs the newest release (the **stable** update
channel). To run the newest, unreleased code from `main` instead (the
**beta** channel), put `LEDMATRIX_CHANNEL=beta` in front of `bash`:
```bash
curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | LEDMATRIX_CHANNEL=beta bash
```
A manual clone starts on `main`; add `--beta` to `first_time_install.sh`
to stay on it, or leave it off and the first update after the next
release moves the device onto releases. You can switch channels later on
the General tab.
**Expected Behavior after install:** **Expected Behavior after install:**
- LED matrix will light up - LED matrix will light up
- A fresh install ships only the bundled `starlark-apps` and - A fresh install ships only the bundled `starlark-apps` and
+250 -24
View File
@@ -4,14 +4,17 @@ The display process serves a Unix socket that the web interface uses to send
it commands and get an answer back. It replaces the cache-file "mailboxes" on it commands and get an answer back. It replaces the cache-file "mailboxes" on
the SD card one command at a time. Stage 1 carries on-demand start, stop and the SD card one command at a time. Stage 1 carries on-demand start, stop and
status. Stage 2 makes those commands land within a frame on every kind of status. Stage 2 makes those commands land within a frame on every kind of
screen, and adds `brightness.set` and `plugin.reload`. The file mailbox stays screen, and adds `brightness.set` and `plugin.reload`. Stage 3 adds a state
as a fallback for one release. stream (`state.get`, `state.subscribe`), so the web interface reads what the
display is doing from the socket instead of from cache files the display
wrote to the SD card. The file mailbox and the cache keys stay as a fallback
for one release.
| | | | | |
|---|---| |---|---|
| Socket | `/run/ledmatrix/control.sock` (tmpfs) | | Socket | `/run/ledmatrix/control.sock` (tmpfs) |
| Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` | | Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` |
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop`, `POST /api/v3/plugins/update` (reload), `POST /api/v3/config/main` (brightness) | | Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop`, `POST /api/v3/plugins/update` (reload), `POST /api/v3/config/main` (brightness); and through [`web_interface/display_state.py`](../web_interface/display_state.py) (the state stream), `GET /api/v3/display/current-status`, `/display/on-demand/status`, `/plugins/installed` (`runtime`), `/plugins/state` and the reconciliations, `/health` (`display_loop`) |
| Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it | | Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it |
| Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it | | Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it |
@@ -36,6 +39,11 @@ running, the socket does not exist, and the web interface knows right away.
## Protocol (version 1) ## Protocol (version 1)
Stage 3 is still version 1: `state.get` and `state.subscribe` are new
commands, and a stage-2 display answers them `unknown_command`, which the
web interface treats as "no socket" and falls back from.
**Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at **Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at
most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with
`ensure_ascii`, so a newline never appears inside a message. A connection `ensure_ascii`, so a newline never appears inside a message. A connection
@@ -75,6 +83,8 @@ one. Clients branch on `error.code`, never on the message text.
| `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly | | `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly |
| `brightness.set` | `{brightness: int 0-100}` | `{brightness, panel_brightness, dimmed, display_active}` | queued, awaited (2 s) | | `brightness.set` | `{brightness: int 0-100}` | `{brightness, panel_brightness, dimmed, display_active}` | queued, awaited (2 s) |
| `plugin.reload` | `{plugin_id}` | `{plugin_id, reloaded: true, version, modes}` | queued, awaited (10 s) | | `plugin.reload` | `{plugin_id}` | `{plugin_id, reloaded: true, version, modes}` | queued, awaited (10 s) |
| `state.get` | `{since?, epoch?}` | a state snapshot (see "The state stream") | answered directly |
| `state.subscribe` | — | a state snapshot, then pushed `state` / `tick` events | answered directly, then a stream |
`duration` is a number of seconds, or a numeric string. `0`, `null` or `""` `duration` is a number of seconds, or a numeric string. `0`, `null` or `""`
mean "until stopped". `pinned` must be a real boolean: the REST route has mean "until stopped". `pinned` must be a real boolean: the REST route has
@@ -130,6 +140,19 @@ web interface treats like any other socket failure and falls back from, and
`hello` lists the commands a display knows. The version changes only when the `hello` lists the commands a display knows. The version changes only when the
envelope or the meaning of an existing command changes. envelope or the meaning of an existing command changes.
**Events.** `state.subscribe` is the one command with more than one message
in reply. After its response, the display pushes events on the same
connection until either side hangs up:
```json
{"v": 1, "id": "<the subscribe id>", "event": "state", "result": {...a state snapshot...}}
{"v": 1, "id": "<the subscribe id>", "event": "tick", "result": {"version": 7, "epoch": "…", "pid": 812, "served_at": 1790000000.1, "changed": false, "loop": {...}, "volatile": {"display": {"last_updated": 1790000000.0}, "...": "..."}}}
```
An event has `event` where a response has `ok`, which is how a reader tells
them apart. The client sends nothing after the subscribe; anything it does
send is ignored.
**Error codes:** `bad_json`, `bad_request`, `message_too_large`, **Error codes:** `bad_json`, `bad_request`, `message_too_large`,
`unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full, `unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full,
or too many connections), `forbidden` (peer credentials refused), `internal`. or too many connections), `forbidden` (peer credentials refused), `internal`.
@@ -147,6 +170,177 @@ print(client.brightness_set(60))
EOF EOF
``` ```
## The state stream (stage 3)
Before stage 3 the web interface learned what the display was doing by
reading files the display kept writing:
| What | Written by the display | How often | Medium |
|---|---|---|---|
| current mode, plugin, `is_display_active`, `on_demand_active` | `display_current_state` | every mode change, every flag change, and every 30 s | cache (SD card) |
| on-demand session | `display_on_demand_state` | on each on-demand event | cache (SD card) |
| plugin runtime snapshot (#690) | `plugin_runtime_snapshot` | on a change (at most every 10 s), else every 60 s | cache (SD card) |
| render-loop liveness (#687) | `display-heartbeat.json` | every 5 s | tmpfs |
Now the display also keeps the same state in memory and serves it on the
socket.
**The snapshot.** `state.get` and `state.subscribe` answer with one object:
```json
{"schema": 1, "version": 42, "epoch": "3f9c0d1e2a4b5c6d", "pid": 812,
"served_at": 1790000000.1, "changed": true,
"loop": {"heartbeat_age_seconds": 1.8, "armed": true, "stale_after": 60.0},
"state": {
"display": {"mode": "nfl_live", "plugin_id": "football-scoreboard", "mode_index": 3,
"total_modes": 9, "on_demand_active": false, "is_display_active": true,
"last_updated": 1790000000.0},
"on_demand": {"active": false, "status": "idle", "...": "as display_on_demand_state"},
"brightness": {"brightness": 80, "panel_brightness": 40, "dimmed": true},
"plugins": {"schema": 1, "running": true, "published_at": 1789999998.5, "...": "as plugin_runtime_snapshot"},
"loop": {"heartbeat_age_seconds": 1.8, "armed": true, "stale_after": 60.0}
}}
```
- `display` and `on_demand` are the dicts the cache keys hold, `plugins` is
the runtime snapshot (`build_runtime_snapshot`), and `brightness` is the
configured level, what the panel shows now, and whether the dim schedule
has it dimmed. A section not published yet is `null`.
- `loop` is not published: the display measures it when it answers, from
the render thread's last beat in memory (`RenderWatchdog.liveness()`),
the same beat that writes the heartbeat file. So it keeps ageing while the
render thread is stuck, and the socket's connection threads still answer.
`heartbeat_age_seconds` is `null` until the loop has drawn its first frame.
- `version` goes up whenever a section changes, ignoring the timestamps that
move on every publish (`last_updated`, `remaining`, `published_at`). It
counts within an `epoch`, one run of the display process, so a reader that
sees a new `epoch` has a restarted display.
- `state.get` with `since` and `epoch` from an earlier answer gets just
`{changed: false, version, epoch, pid, served_at, loop, volatile}` while
nothing has changed. `volatile` is `{section: {key: value}}`: the current
values of those ignored timestamps, which the reader merges into the copy
it has. They don't make a new version, but they are still news:
`display.last_updated` is how a reader knows the render thread is still
publishing, and `plugins.published_at` the runtime publisher. Without
them a reader's copy kept the timestamps of the last real change, so a
mode on screen for over 120 s read as unknown.
- A snapshot that would not fit in a message (hundreds of plugins) is sent
without `plugins`, and `truncated: ["plugins"]` says so. Readers then use
the cache for that section only.
**The stream.** `state.subscribe` answers with the snapshot, then:
- a `state` event (a full snapshot) whenever the version changes, and
- a `tick` at least every 5 s (`SUBSCRIBE_KEEPALIVE_SECONDS`) when nothing
changed. It is the short `changed: false` answer, so it carries `loop`
(a stalled render loop shows up within one tick) and `volatile` (the
timestamps stay as fresh as the writers keep them), and it tells the
reader the connection is alive.
A slow reader is never sent a backlog: each event is the latest version, so
one that falls behind skips the versions in between. A reader that has heard
nothing for 15 s (three keepalives) stops trusting its copy.
**Who publishes, and when.** All of it is in memory, with no disk writes:
- the render thread, at the places it already published the cache keys:
`display` and `brightness` on every pass of
`_publish_current_mode_state_if_changed()` (every loop pass, and every
`_service_pending_changes()` in a dwell, a scrolling screen or Vegas), and
`on_demand` in `_publish_on_demand_state()`. Every pass refreshes
`display.last_updated`, so a reader can tell when the render thread has
stopped publishing, just as the cache key's 120 s `max_age` does.
- the plugin runtime publisher's thread, on every 5 s tick: the snapshot is
rebuilt when the state machine changed, otherwise only its `published_at`
moves. A change reaches subscribers within a tick, without the cache's
10 s throttle.
Publishing is a hand-off, as the command queue is in the other direction.
The hub (`StateHub` in [`src/ipc/server.py`](../src/ipc/server.py)) holds a
lock only to swap a dict reference, compare it with the last one and bump the
version. Every socket write happens on the subscriber's own connection
thread. The render thread never waits for a reader.
### Readers in the web interface
[`web_interface/display_state.py`](../web_interface/display_state.py) holds
one `state.subscribe` connection per web process
(`src.ipc.client.StateSubscription`, a daemon thread, started on the first
read and reconnecting with a backoff of 1 s up to 30 s). A route answers
from the latest pushed snapshot in memory. Before the subscription has one,
the route asks once with `state.get` (0.5 s timeout). When neither works, it
reads the cache keys and the heartbeat file as before:
| Route | From the socket | Fallback |
|---|---|---|
| `GET /api/v3/display/current-status` | `state.display` | `display_current_state` |
| `GET /api/v3/display/on-demand/status` | `state.on_demand`, with `remaining` worked out from `expires_at` now | `display_on_demand_state` |
| `GET /api/v3/plugins/installed` (`runtime`), `/plugins/state`, `POST /plugins/state/reconcile` and the startup reconciliation | `state.plugins` + `state.loop` | `plugin_runtime_snapshot` + `display-heartbeat.json` |
| `GET /api/v3/health` (`checks.display_loop`) | `state.loop` | `display-heartbeat.json` |
Each answer says where it came from: `source: "socket" | "cache"` (or
`"heartbeat_file"` for the health check).
The SSE display stream (`/api/v3/stream/display`) reads the preview frame
file, not a cache key, so it does not change.
**The same verdicts either way.** The socket's answers are judged by the
rules the cache readers apply (#726):
- the runtime view is `stalled` when the render loop's heartbeat age is at
least `HEARTBEAT_STALE_SECONDS` (60 s, the health check's threshold), and
then reports no per-plugin facts;
- it is `stale` when the snapshot is older than its `stale_after` (the
publisher thread stopped);
- with no beat yet, the snapshot is judged on its own;
- there is no pid check, because the display that answered is alive;
- a `display` section the render thread has not refreshed for 120 s reads
as unknown, as the cache key does once it ages out.
The age a reader uses is the age the display measured, plus the time since
the snapshot arrived.
### Fewer SD writes
The cache keys are still written, for one release, as the fallback. While
the socket serves the readers, the display writes two of them less often.
"Serves the readers" means a subscriber is connected, or a `state.get` came
within the last 60 s (`StateHub.readers_active()`):
- `display_current_state` is no longer written on every mode change: once
every 60 s (`CURRENT_STATE_RELAXED_REFRESH_SECONDS`, inside the readers'
120 s `max_age`), and at once when `is_display_active` or
`on_demand_active` changes.
- `plugin_runtime_snapshot`'s refresh goes from 60 s to 120 s
(`RELAXED_REFRESH_INTERVAL`), and the snapshot says so in its own
`refresh_interval` and `stale_after` (360 s). Changes are still written at
once, at most every 10 s.
`display_on_demand_state` is written only on events, so it is unchanged.
The heartbeat file is on tmpfs, so it costs no SD writes, and it stays: the
automatic update's health check reads it.
This is safe because the relaxed rate only applies while readers are using
the socket. If they stop (the web interface loses the socket, or is stopped),
the next publish after the reader window writes a changed mode at once, and
the runtime refresh goes back to 60 s. A fallback reader in that window sees
a mode up to 60 s old, never one older than its `max_age`.
Measured with fake clocks (`test_cache_writes_per_minute_with_and_without_socket_readers`
in `test/test_state_stream_readers.py`), for a rotation of 15 s screens:
| Key | Writes/min, no socket readers | Writes/min, socket readers |
|---|---|---|
| `display_current_state` | 4.0 | 1.0 |
| `plugin_runtime_snapshot` | 1.0 | 0.5 |
| Total | 5.0 | 1.5 |
That is 70% fewer writes for these keys: about 2,200 a day instead of 7,200.
Shorter screens save more, because the old rate followed the mode changes.
A display that rarely changes mode (one plugin, a long live game) saves less. Plugin
data caches, the error snapshot and font usage are written by other code
and are not affected.
## How the display applies a command ## How the display applies a command
The server's threads never touch rendering. A connection thread parses the The server's threads never touch rendering. A connection thread parses the
@@ -279,7 +473,16 @@ block the render loop or crash it:
process created. process created.
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind - **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
as before. The web interface then uses the mailbox. as before. The web interface then uses the mailbox, and reads the cache
keys and the heartbeat file.
- **Subscribers (stage 3).** A `state.subscribe` connection gives its request
slot back and takes one of 4 subscriber slots (`MAX_SUBSCRIBERS`). A fifth
gets `busy`. So a few browsers' web processes holding streams can never
use up the 8 slots that commands need. Each subscriber has its own thread.
A send that cannot finish within the 2 s IO timeout (a reader that stopped
reading) drops that subscriber. Nothing else waits for it, and the render
thread only publishes to the hub. `close()` wakes every subscriber, so
they end at once.
## Security model ## Security model
@@ -313,8 +516,10 @@ read its state, set the brightness, and reload a plugin the display is
already running, all of which anyone who can reach the web UI can already do already running, all of which anyone who can reach the web UI can already do
(the last by restarting the display). Nothing on the socket runs a shell, (the last by restarting the display). Nothing on the socket runs a shell,
writes a file, or names a path, and `plugin.reload` cannot make the display writes a file, or names a path, and `plugin.reload` cannot make the display
import a plugin it was not running. Stage 2 changed none of the access rules import a plugin it was not running. Stages 2 and 3 changed none of the
above. access rules above. The state stream carries what the cache keys already
held, and those are readable by the same group. A subscriber goes through
the same connect-time and peer-credential checks as any other connection.
**Development.** A display that is not root and cannot write to **Development.** A display that is not root and cannot write to
`/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the `/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the
@@ -352,29 +557,34 @@ device never touches the live display.
"which sections changed" ack had no reader: the web interface knows "which sections changed" ack had no reader: the web interface knows
what it saved. A reload from the socket thread would also run every what it saved. A reload from the socket thread would also run every
config subscriber on a second thread beside the watcher's. config subscriber on a second thread beside the watcher's.
3. **A state stream.** A `subscribe` command that keeps the connection open 3. **A state stream (done).** `state.get` (a versioned snapshot) and
and pushes events: mode changes, on-demand state (including the outcome of `state.subscribe` (the snapshot, then pushed changes and keepalive ticks)
an acked on-demand command, which today is only published), plugin carry the current mode, the on-demand state (including the outcome of an
runtime state, the outcome of a reload that answered `pending`, and the acked on-demand command), the brightness, the plugin runtime snapshot and
heartbeat. It replaces the polled `display_current_state`, the render loop's liveness, all served from memory (see "The state
`plugin_runtime_snapshot` (#690) and `display-heartbeat.json` (#687) for stream"). The web interface's readers use it and fall back to the cache
readers that hold a connection. The web interface relays it to its keys and the heartbeat file. `display_current_state` and
existing SSE stream. The files remain for one release for older readers. `plugin_runtime_snapshot` are written less often while it serves them.
- The server's per-connection threads (8 at most) do not suit long-lived The keys remain for one release.
subscribers. A subscriber needs its own bound and a writer that drops - Left for later: the outcome of a `plugin.reload` that answered
events for a slow reader rather than blocking the display. `pending` is visible only as the plugin's new `loaded_version` in
- Events are produced on the render thread, so publishing must be a `state.plugins`, not as an event of its own.
non-blocking hand-off, like the queue in the other direction. - Left for later: the SSE display stream reads the preview frame, not
- The store's install of an already-enabled plugin, and an uninstall that state, so nothing relays the stream to the browser yet. A browser still
keeps its config, still answer `restart_required`. With the stream they polls the REST routes, which now answer from memory.
can use a load/unload command and report the result the same way the - Left for later: the store's install of an already-enabled plugin, and
update route does now. an uninstall that keeps its config, still answer `restart_required`.
They can now use a load/unload command and report the result the same
way the update route does.
4. **Retire the mailboxes.** After a release in which every device has had the 4. **Retire the mailboxes.** After a release in which every device has had the
socket, the web interface stops writing `display_on_demand_request`, and socket, the web interface stops writing `display_on_demand_request`, and
the display stops polling it, logging the plugins that still write it so the display stops polling it, logging the plugins that still write it so
they can move to an in-process `request_display()`. The other cache keys they can move to an in-process `request_display()`. The other cache keys
used as messages (`plugin_error_clear_request` and the remaining used as messages (`plugin_error_clear_request` and the remaining
`display_*` keys) move to the socket or to tmpfs. `display_*` keys) move to the socket or to tmpfs. The display also stops
writing `display_current_state`, `display_on_demand_state` and
`plugin_runtime_snapshot` once the web interface no longer falls back to
them.
## Checking it on a device ## Checking it on a device
@@ -405,3 +615,19 @@ sudo journalctl -u ledmatrix | grep -E "Brightness set|Reload(ing|ed) plugin"
`unknown_command` in `brightness_socket_error` or `reload_error` means the `unknown_command` in `brightness_socket_error` or `reload_error` means the
display runs a stage-1 build: restart it once to pick up this one. display runs a stage-1 build: restart it once to pick up this one.
The state stream:
```bash
curl -s localhost:5000/api/v3/display/current-status # ... "source": "socket"
curl -s localhost:5000/api/v3/health | python3 -m json.tool | grep -A3 display_loop
python3 - <<'EOF'
from src.ipc import client # run from the project directory
snap = client.state_get()
print(snap['version'], snap['epoch'], snap['loop'], snap['state']['display'])
EOF
```
`"source": "cache"` means the web interface could not use the socket: the
display is stopped, predates stage 3, or the web user is not in the
socket's group.
+20 -1
View File
@@ -33,6 +33,9 @@ in again (services pick them up on restart).
| `/run/ledmatrix/control.sock` | `root` : cache directory's group (`ledmatrix`) | `660` | The display's control socket; only root and that group can connect. See [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md#security-model) | | `/run/ledmatrix/control.sock` | `root` : cache directory's group (`ledmatrix`) | `660` | The display's control socket; only root and that group can connect. See [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md#security-model) |
| `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them | | `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them |
| `/etc/sudoers.d/ledmatrix_web`, `ledmatrix_wifi` | `root` | `440` | | | `/etc/sudoers.d/ledmatrix_web`, `ledmatrix_wifi` | `root` | `440` | |
| `/usr/local/sbin/ledmatrix-refresh-units` | `root:root` | `755` | Copy of `scripts/install/ledmatrix_refresh_units.py`, installed by `install_service.sh`. Outside the project so the web user cannot edit what sudo runs |
| `/var/lib/ledmatrix/unit-backup/` | `root` | `700` | The units the last refresh replaced, for the automatic update's rollback |
| `/etc/systemd/system/ledmatrix*.service`, `.path` | `root:root` | `644` | Readable so the web interface can compare them with the templates after an update |
What keeps it that way at runtime: What keeps it that way at runtime:
@@ -86,6 +89,20 @@ password:
- `journalctl -u ledmatrix.service *`, `-u ledmatrix *`, `-t ledmatrix *`, - `journalctl -u ledmatrix.service *`, `-u ledmatrix *`, `-t ledmatrix *`,
tagged `NOEXEC`: journalctl opens a pager on a terminal, and a shell tagged `NOEXEC`: journalctl opens a pager on a terminal, and a shell
escape from that pager would be a root shell escape from that pager would be a root shell
- `/usr/local/sbin/ledmatrix-refresh-units ""` and
`/usr/local/sbin/ledmatrix-refresh-units --restore` — exactly these two
command lines (`""` means "no arguments"). After an update the first
installs the systemd units whose templates changed and runs
`systemctl daemon-reload`; the automatic update's rollback runs the second
to put the previous units back. The helper takes nothing from the caller:
the project folder and the web user come from the installed, root-owned
`ledmatrix.service` and `ledmatrix-web.service`. It only replaces the four
units `install_service.sh` installs, only if they are already installed,
and refuses a template that would change a unit's `User=` (root for the
display, the web user for the rest) or `WorkingDirectory=`, or that is a
symlink, not a regular file, or over 64 KB. It grants nothing new: the
templates are files the web user can edit, but so is `run.py`, which the
display service already runs as root.
### `/etc/sudoers.d/ledmatrix_wifi` ### `/etc/sudoers.d/ledmatrix_wifi`
@@ -136,7 +153,9 @@ directory.
| `safe_plugin_rm.sh`, `safe_pip_install.sh` | — | Called by the web interface through sudo | Not for manual use | | `safe_plugin_rm.sh`, `safe_pip_install.sh` | — | Called by the web interface through sudo | Not for manual use |
To reinstall the sudoers rules, run To reinstall the sudoers rules, run
`./scripts/install/configure_web_sudo.sh` (web rules) or `./scripts/install/configure_web_sudo.sh` (web rules; the
`ledmatrix-refresh-units` rules also need the helper itself, which
`sudo ./scripts/install/install_service.sh` installs) or
`./scripts/install/configure_wifi_permissions.sh` (WiFi rules and polkit) as `./scripts/install/configure_wifi_permissions.sh` (WiFi rules and polkit) as
the web user, not with `sudo`. the web user, not with `sudo`.
+25 -6
View File
@@ -331,12 +331,19 @@ by the display process (stale after 120 seconds).
"data": { "data": {
"mode": "nfl_live", "mode": "nfl_live",
"plugin_id": "football-scoreboard", "plugin_id": "football-scoreboard",
"last_updated": 1234567890.123 "last_updated": 1234567890.123,
"source": "socket"
} }
} }
``` ```
When nothing has been published, every field is `null`. When nothing has been published, every field is `null`. `source` is
`socket` when the answer came from the display's state stream over the
control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), and `cache`
when it came from the `display_current_state` cache key (no socket: the
display is stopped or older, or this is Windows). A display whose render
loop has not refreshed its state for 120 seconds is reported with every
field `null`, either way.
### List Display Modes ### List Display Modes
@@ -413,11 +420,16 @@ Get the current on-demand display state.
"returncode": 0, "returncode": 0,
"stdout": "active", "stdout": "active",
"stderr": "" "stderr": ""
} },
"source": "socket"
} }
} }
``` ```
`source` is `socket` (the display's state stream, with `remaining` worked
out at the time of the request) or `cache` (the `display_on_demand_state`
cache key).
With no on-demand request, `state` is With no on-demand request, `state` is
`{"active": false, "status": "idle", "last_updated": null}`. `{"active": false, "status": "idle", "last_updated": null}`.
@@ -552,7 +564,8 @@ List all installed plugins with their status and metadata.
"published_at": 1790000030.0, "published_at": 1790000030.0,
"age_seconds": 12.4, "age_seconds": 12.4,
"stale_after": 180.0, "stale_after": 180.0,
"heartbeat_age_seconds": 2.1 "heartbeat_age_seconds": 2.1,
"source": "socket"
} }
} }
} }
@@ -581,7 +594,10 @@ the display is hung or died), `stopped` (the display shut down) or `unknown`
(nothing published yet). Unless it is `live`, every one of those fields is (nothing published yet). Unless it is `live`, every one of those fields is
`null`. `heartbeat_age_seconds` is the heartbeat's age when it was taken into `null`. `heartbeat_age_seconds` is the heartbeat's age when it was taken into
account, `null` otherwise (no heartbeat, as on the dev server, or one from account, `null` otherwise (no heartbeat, as on the dev server, or one from
another process). Health and metrics are at [`/plugins/health`](#get-plugin-health) another process). `runtime.source` is `socket` when the snapshot and the
heartbeat age came from the display's state stream over the control socket,
and `cache` when they came from the `plugin_runtime_snapshot` cache key and
the heartbeat file; the rules above are the same for both. Health and metrics are at [`/plugins/health`](#get-plugin-health)
and `/plugins/metrics`. and `/plugins/metrics`.
`vegas_participation` is what Vegas mode does with the plugin: `"scroll"`, `vegas_participation` is what Vegas mode does with the plugin: `"scroll"`,
@@ -2265,7 +2281,10 @@ display snapshot. `data.status` is `healthy` or `degraded`, with
(with `heartbeat_age_seconds`), `stalled` (no heartbeat for 60s: the panel is (with `heartbeat_age_seconds`), `stalled` (no heartbeat for 60s: the panel is
frozen even if the service is active; the status turns `degraded`), or frozen even if the service is active; the status turns `degraded`), or
`not_reported` when the display writes none (not started yet, the dev server, `not_reported` when the display writes none (not started yet, the dev server,
Windows), which does not affect the status. Windows), which does not affect the status. Its `source` is `socket` when the
age came from the display's state stream over the control socket (measured
in memory by the display) and `heartbeat_file` when it came from
`/run/ledmatrix/display-heartbeat.json`.
Open even when the web login is on, for uptime monitors; a caller that is not Open even when the web login is on, for uptime monitors; a caller that is not
logged in (and has no token) then gets only `{"status": "success", "data": logged in (and has no token) then gets only `{"status": "success", "data":
+14 -2
View File
@@ -246,8 +246,16 @@ mean exactly 10 ms, so a ticker stalling on half its frames still averages to a
healthy 100 fps. The stats line reports the tail for that reason — read the healthy 100 fps. The stats line reports the tail for that reason — read the
percentiles, not the fps. percentiles, not the fps.
Every scroller emits one line every 5 seconds covering *every* frame in that Every scroller summarises each 5-second window, covering *every* frame in it,
window, tagged with the plugin it came from: in one line tagged with the plugin it came from. At the default log level the
line reaches the journal only when it is worth reading: a **degraded** window
(frame rate below 90% of the rate the window was locked to, i.e. 1 / its own
median -- the same 0.9 Vegas's `Vegas FPS` line uses -- or more than 1% of its
frames stalled), the first window after one (the recovery), and otherwise once
every 5 minutes per scroller as a heartbeat, so silence means stopped rather
than fine. Every window is logged at DEBUG: to see them all, run the display
with `-d` or `LEDMATRIX_DEBUG=true` (see
[CONFIG_DEBUGGING.md](CONFIG_DEBUGGING.md#enable-debug-logging)).
```bash ```bash
journalctl -u ledmatrix --since "-10min" --no-pager | grep "Scroll frame stats" journalctl -u ledmatrix --since "-10min" --no-pager | grep "Scroll frame stats"
@@ -286,6 +294,10 @@ journalctl -u ledmatrix --since "-3h" --no-pager | grep "Scroll frame stats" \
| sort -k7 -rn | sort -k7 -rn
``` ```
At the default log level that ranks the windows the journal kept -- the
degraded ones, recoveries and heartbeats -- so it over-weights bad windows;
rank a debug run for an unbiased average, or soak the rig (below).
The `$2 < 1000` guard drops windows whose median is a whole second or more. The `$2 < 1000` guard drops windows whose median is a whole second or more.
Those are not frames. Until the idle-gap fix in `log_frame_rate()`, the first Those are not frames. Until the idle-gap fix in `log_frame_rate()`, the first
frame of every scroll was timed against the end of the *previous* scroll, so frame of every scroll was timed against the end of the *previous* scroll, so
+91 -3
View File
@@ -84,6 +84,43 @@ python3 web_interface/start.py
### Installation & Build Issues ### Installation & Build Issues
#### "This version of Raspberry Pi OS is not supported"
LEDMatrix installs on Raspberry Pi OS Lite **Trixie** (Debian 13, Python
3.13) or **Bookworm** (Debian 12, Python 3.11). The installer checks
`/etc/os-release` before it changes anything and stops on anything else.
**Check what you have:**
```bash
grep -E '^(PRETTY_NAME|VERSION_ID)=' /etc/os-release
python3 --version
```
**Solutions:**
- `VERSION_ID="11"` (Bullseye) or older: flash a new card with Raspberry Pi
Imager, choosing Raspberry Pi OS Lite (64-bit). Trixie is recommended;
Bookworm (Legacy) also works. An in-place upgrade from Bullseye is not
supported by Raspberry Pi and is not worth the risk.
- "Desktop environment detected": use the Lite image, not the desktop one.
- "python3 is Python 3.x; LEDMatrix needs Python 3.11 or newer": something
has replaced the system `python3`. Point it back at the OS's own Python
(`/usr/bin/python3` should be 3.11 on Bookworm, 3.13 on Trixie).
- `sudo bash scripts/check_system_compatibility.sh` runs the same checks
without installing anything.
#### "This Pi manages its network with dhcpcd, not NetworkManager"
A warning, not an error: the install carries on and the display works. But
choosing a WiFi network from the web page and the `LEDMatrix-Setup` hotspot
both need NetworkManager, the default on Bookworm and Trixie. It appears
when dhcpcd was selected in `raspi-config`. Switch back with a keyboard and
screen attached (or over Ethernet), since the WiFi connection drops briefly:
```bash
sudo raspi-config # Advanced Options -> Network Config -> NetworkManager
sudo reboot
```
#### Step 6 fails: "Failed building wheel for rgbmatrix" #### Step 6 fails: "Failed building wheel for rgbmatrix"
**Symptoms:** **Symptoms:**
@@ -328,6 +365,49 @@ commit, then switches to releases on its own.
3. **Local changes after a channel switch:** edits that no longer fit the new 3. **Local changes after a channel switch:** edits that no longer fit the new
version are kept in the git stash rather than lost; `git stash list` version are kept in the git stash rather than lost; `git stash list`
shows them as "LEDMatrix autostash before update". shows them as "LEDMatrix autostash before update".
4. **A new install is on a release, not `main`.** The one-shot installer
checks out the newest release. For the newest code instead, install with
`LEDMATRIX_CHANNEL=beta`:
```bash
curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | LEDMATRIX_CHANNEL=beta bash
```
---
#### Issue: "service settings ... are not applied yet" after an update
**Symptoms:**
- Update Code's message, or the web interface log, says an update changes
service settings that are not applied yet, and to run the installer
- The display logs `ledmatrix.service differs from systemd/ledmatrix.service`
at startup
**Explanation:** updates install the systemd units a new version changes
through the root helper `/usr/local/sbin/ledmatrix-refresh-units`, which the
installer sets up and grants to the web user in
`/etc/sudoers.d/ledmatrix_web`. A device installed before that has neither,
so the new unit settings (for example the display's watchdog) wait for a
reinstall. The update itself is fine.
**Solution:** re-run the installer once, as root:
```bash
cd ~/LEDMatrix
sudo ./first_time_install.sh
# or, lighter: install the units and helper, then the sudo rules
sudo ./scripts/install/install_service.sh
./scripts/install/configure_web_sudo.sh
```
Check it worked:
```bash
ls -l /usr/local/sbin/ledmatrix-refresh-units # root root, rwxr-xr-x
sudo -l | grep ledmatrix-refresh-units # the two rules
```
A message that the helper **refused** a unit (`refusing to install it`)
means a template in `systemd/` was edited so that it would run as another
account or from another folder. The message names the template. Look at
what changed with `git diff -- systemd/`, save any edit you want to keep,
then restore only that file, for example
`git checkout -- systemd/ledmatrix-web.service`.
--- ---
@@ -364,9 +444,16 @@ commit, then switches to releases on its own.
5. **Check required services:** 5. **Check required services:**
```bash ```bash
systemctl is-active NetworkManager # must say "active"
sudo systemctl status hostapd sudo systemctl status hostapd
sudo systemctl status dnsmasq sudo systemctl status dnsmasq
``` ```
On a fresh install `hostapd` shows as **masked**. That is expected, on
Bookworm and Trixie alike: Debian's hostapd package masks the service
when it is installed without a configuration, so the hotspot is brought
up through NetworkManager instead (look for `nmcli hotspot fallback` in
`journalctl -u ledmatrix-wifi-monitor`). If NetworkManager is not
active, see "This Pi manages its network with dhcpcd" above.
6. **Manually enable AP mode:** 6. **Manually enable AP mode:**
```bash ```bash
@@ -590,9 +677,10 @@ stack into the log, so it says which plugin was stuck.
apart, so a plugin that hangs on every start does not restart the display apart, so a plugin that hangs on every start does not restart the display
hundreds of times an hour. hundreds of times an hour.
4. **Is the watchdog installed?** Installs from before it keep their old unit 4. **Is the watchdog installed?** Updates install new unit settings once the
until the installer is re-run (a startup warning says the unit differs installer has set up `ledmatrix-refresh-units`; installs from before that
from its template): keep their old unit until the installer is re-run (a startup warning says
the unit differs from its template):
```bash ```bash
systemctl show -p WatchdogUSec ledmatrix # 2min once running; 0 = not installed systemctl show -p WatchdogUSec ledmatrix # 2min once running; 0 = not installed
sudo ./scripts/install/install_service.sh sudo ./scripts/install/install_service.sh
+157 -44
View File
@@ -47,38 +47,51 @@ if echo "${DEVICE_MODEL:-}" | grep -qi "Raspberry Pi 5"; then
echo "Raspberry Pi 5 detected — will verify RP1 library support." echo "Raspberry Pi 5 detected — will verify RP1 library support."
fi fi
# Check OS version - must be Raspberry Pi OS Lite (Trixie) # Check OS version - must be Raspberry Pi OS Lite, Bookworm or Trixie.
# The rules live in scripts/install/lib_os.sh, shared with
# scripts/check_system_compatibility.sh.
echo "" echo ""
echo "Checking operating system requirements..." echo "Checking operating system requirements..."
echo "----------------------------------------" echo "----------------------------------------"
OS_CHECK_FAILED=0 OS_CHECK_FAILED=0
OS_RELEASE=""
if [ -f /etc/os-release ]; then OS_LIB="$(cd "$(dirname "$0")" && pwd)/scripts/install/lib_os.sh"
. /etc/os-release if [ ! -f "$OS_LIB" ]; then
echo "Detected OS: $PRETTY_NAME" echo "✗ ERROR: $OS_LIB is missing, so the operating system cannot be checked."
echo "Version ID: ${VERSION_ID:-unknown}" echo " Your LEDMatrix download is incomplete. Download it again and re-run this script:"
echo " git clone https://github.com/ChuckBuilds/LEDMatrix.git"
exit 1
fi
# shellcheck source=scripts/install/lib_os.sh
. "$OS_LIB"
# Check if it's Raspberry Pi OS or Debian if [ -r "$LM_OS_RELEASE_FILE" ]; then
if [[ "$ID" != "raspbian" ]] && [[ "$ID" != "debian" ]]; then echo "Detected OS: $(lm_os_field PRETTY_NAME)"
echo "✗ ERROR: This script requires Raspberry Pi OS (raspbian/debian)" OS_VERSION_ID=$(lm_os_field VERSION_ID)
echo " Detected OS ID: $ID" echo "Version ID: ${OS_VERSION_ID:-unknown}"
OS_CHECK_FAILED=1
fi
# Check if it's Debian 13 (Trixie) if OS_RELEASE=$(lm_os_release); then
if [ "${VERSION_ID:-0}" != "13" ]; then echo "✓ $(lm_release_label "$OS_RELEASE") detected"
echo "✗ ERROR: This script requires Raspberry Pi OS Lite (Trixie) - Debian 13"
echo " Detected version: ${VERSION_ID:-unknown}"
echo " Please upgrade to Raspberry Pi OS Lite (Trixie) before continuing"
OS_CHECK_FAILED=1
else else
echo "✓ Debian 13 (Trixie) detected" OS_ID=$(lm_os_field ID)
if [[ "$OS_ID" != "raspbian" ]] && [[ "$OS_ID" != "debian" ]]; then
echo "✗ ERROR: This script requires Raspberry Pi OS (raspbian/debian)"
echo " Detected OS ID: ${OS_ID:-unknown}"
else
echo "✗ ERROR: This version of Raspberry Pi OS is not supported"
echo " Detected version: ${OS_VERSION_ID:-unknown}"
echo " Supported: Trixie (Debian 13) and Bookworm (Debian 12)"
fi
OS_CHECK_FAILED=1
fi fi
# Check if it's the Lite version (no desktop environment) # Check if it's the Lite version (no desktop environment)
# Check for desktop packages or desktop services # Check for desktop packages or desktop services
DESKTOP_DETECTED=0 DESKTOP_DETECTED=0
if dpkg -l | grep -qE "^ii.*raspberrypi-ui-mods|^ii.*lxde|^ii.*xfce|^ii.*gnome|^ii.*kde"; then # grep without -q: -q exits at the first match, dpkg then dies of SIGPIPE,
# and pipefail turns a found desktop into "not found".
if dpkg -l | grep -E "^ii.*raspberrypi-ui-mods|^ii.*lxde|^ii.*xfce|^ii.*gnome|^ii.*kde" >/dev/null; then
DESKTOP_DETECTED=1 DESKTOP_DETECTED=1
fi fi
if systemctl list-units --type=service --state=running 2>/dev/null | grep -qE "lightdm|gdm3|sddm|lxdm"; then if systemctl list-units --type=service --state=running 2>/dev/null | grep -qE "lightdm|gdm3|sddm|lxdm"; then
@@ -96,23 +109,52 @@ if [ -f /etc/os-release ]; then
echo "✓ Lite version confirmed (no desktop environment)" echo "✓ Lite version confirmed (no desktop environment)"
fi fi
else else
echo "✗ ERROR: Could not detect OS version (/etc/os-release not found)" echo "✗ ERROR: Could not detect OS version ($LM_OS_RELEASE_FILE not found)"
OS_CHECK_FAILED=1 OS_CHECK_FAILED=1
fi fi
# Python: whatever python3 the release ships (3.11 on Bookworm, 3.13 on
# Trixie). Checked only when python3 is already there -- Step 1 installs it
# otherwise, and on a supported release that brings the release's own version.
if [ "$OS_CHECK_FAILED" -eq 0 ]; then
if PYTHON3_VERSION=$(lm_python_version); then
case "$(lm_python_check "$PYTHON3_VERSION")" in
ok)
echo "✓ Python $PYTHON3_VERSION detected"
;;
too-old)
echo "✗ ERROR: python3 is Python $PYTHON3_VERSION; LEDMatrix needs Python 3.$LM_PYTHON_MIN_MINOR or newer"
echo " $(lm_release_label "$OS_RELEASE") ships Python $(lm_release_python "$OS_RELEASE"). Something on this"
echo " system has changed which Python 'python3' runs; point it back at the system Python."
OS_CHECK_FAILED=1
;;
*)
echo "⚠ python3 is Python $PYTHON3_VERSION, which LEDMatrix has not been tested with"
echo " (tested: 3.$LM_PYTHON_MIN_MINOR to 3.$LM_PYTHON_MAX_MINOR). Continuing anyway."
;;
esac
else
echo "python3 not found yet; Step 1 installs it."
fi
fi
if [ "$OS_CHECK_FAILED" -eq 1 ]; then if [ "$OS_CHECK_FAILED" -eq 1 ]; then
echo "" echo ""
echo "Installation cannot continue. Please install Raspberry Pi OS Lite (Trixie) and try again." echo "Installation cannot continue."
echo "" lm_print_supported_os_help
echo "To install Raspberry Pi OS Lite (Trixie):"
echo " 1. Download from: https://www.raspberrypi.com/software/operating-systems/"
echo " 2. Select 'Raspberry Pi OS Lite (64-bit)' with Debian 13 (Trixie)"
echo " 3. Flash to SD card using Raspberry Pi Imager"
echo " 4. Boot and run this script again"
exit 1 exit 1
fi fi
echo "✓ OS requirements met" echo "✓ OS requirements met"
# WiFi setup (the web page's WiFi tab and the LEDMatrix-Setup hotspot) needs
# NetworkManager. Both releases use it by default; say so plainly if this Pi
# does not, but carry on -- the display itself does not depend on it.
case "$(lm_network_stack)" in
networkmanager) echo "✓ NetworkManager is managing the network" ;;
dhcpcd) lm_print_dhcpcd_advice ;;
*) echo "⚠ Could not tell which service manages the network; WiFi setup from the web page needs NetworkManager" ;;
esac
echo "" echo ""
# The user who ran the installer: SUDO_USER once we are running under sudo # The user who ran the installer: SUDO_USER once we are running under sudo
@@ -190,6 +232,51 @@ _sync_rgb_submodule() {
fi fi
return 0 return 0
} }
# LEDMatrix's own changes to the library live in patches/rpi-rgb-led-matrix/ and
# are applied only for the build: _apply_rgb_patches before it, _revert_rgb_patches
# after it, success or not. The checkout is left exactly as it was, so `git pull`
# and _sync_rgb_submodule never meet local modifications in the submodule.
# A patch that no longer applies (a submodule bump, a hand-edited checkout) is
# reported and skipped -- the unpatched library still builds and works, so it is
# never fatal. One that is already applied is left alone and not reverted.
_RGB_APPLIED_PATCHES=()
_apply_rgb_patches() {
local sub="$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master"
local dir="$PROJECT_ROOT_DIR/patches/rpi-rgb-led-matrix" patch name
_RGB_APPLIED_PATCHES=()
[ -d "$dir" ] || return 0
for patch in "$dir"/*.patch; do
[ -f "$patch" ] || continue
name=$(basename "$patch")
if _git_as_repo_owner -C "$sub" apply --check "$patch" >/dev/null 2>&1; then
if _git_as_repo_owner -C "$sub" apply "$patch"; then
_RGB_APPLIED_PATCHES+=("$patch")
echo "Applied library patch $name"
else
echo "⚠ Could not apply library patch $name; building without it"
fi
elif _git_as_repo_owner -C "$sub" apply --reverse --check "$patch" >/dev/null 2>&1; then
echo "Library patch $name is already applied"
else
echo "⚠ Library patch $name does not apply to this checkout; building without it"
fi
done
return 0
}
_revert_rgb_patches() {
local sub="$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master" i
# Last applied first, in case two patches touch the same file.
for ((i = ${#_RGB_APPLIED_PATCHES[@]} - 1; i >= 0; i--)); do
if ! _git_as_repo_owner -C "$sub" apply --reverse "${_RGB_APPLIED_PATCHES[i]}"; then
echo "⚠ Could not revert $(basename "${_RGB_APPLIED_PATCHES[i]}"); restore the checkout with: git -C $sub checkout -- ."
fi
done
_RGB_APPLIED_PATCHES=()
return 0
}
# --- end rpi-rgb-led-matrix checkout helpers --------------------------------- # --- end rpi-rgb-led-matrix checkout helpers ---------------------------------
# Determine the Project Root Directory (where this script is located) # Determine the Project Root Directory (where this script is located)
@@ -223,6 +310,8 @@ SKIP_SWAP=${LEDMATRIX_SKIP_SWAP:-0}
BUILD_JOBS_OVERRIDE=${LEDMATRIX_BUILD_JOBS:-} BUILD_JOBS_OVERRIDE=${LEDMATRIX_BUILD_JOBS:-}
# Weekly automatic updates: 1 on, 0 off, empty = ask (interactive) or leave as is. # Weekly automatic updates: 1 on, 0 off, empty = ask (interactive) or leave as is.
AUTO_UPDATE=${LEDMATRIX_AUTO_UPDATE:-} AUTO_UPDATE=${LEDMATRIX_AUTO_UPDATE:-}
# Update channel written to config.json: stable, beta, or empty = leave as is.
UPDATE_CHANNEL=$(printf '%s' "${LEDMATRIX_CHANNEL:-}" | tr '[:upper:]' '[:lower:]')
usage() { usage() {
cat <<USAGE cat <<USAGE
@@ -240,12 +329,18 @@ Options:
--enable-auto-update Turn on weekly automatic updates (with health --enable-auto-update Turn on weekly automatic updates (with health
check and automatic rollback) check and automatic rollback)
--no-auto-update Leave weekly automatic updates off --no-auto-update Leave weekly automatic updates off
--beta Follow main, the newest code (the beta update
channel). Without it, updates follow releases
(stable). It sets the channel; it does not move
this checkout -- the one-shot installer picks the
version, and so does the next update.
-h, --help Show this help message and exit -h, --help Show this help message and exit
Environment variables (same effect as flags): Environment variables (same effect as flags):
LEDMATRIX_ASSUME_YES=1, RPI_RGB_FORCE_REBUILD=1, LEDMATRIX_SKIP_SOUND=1, LEDMATRIX_ASSUME_YES=1, RPI_RGB_FORCE_REBUILD=1, LEDMATRIX_SKIP_SOUND=1,
LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1, LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1,
LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N, LEDMATRIX_AUTO_UPDATE=1|0 LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N, LEDMATRIX_AUTO_UPDATE=1|0,
LEDMATRIX_CHANNEL=stable|beta
Low-memory devices: Low-memory devices:
On a Pi with under 2GB of RAM the C++ build is limited to fewer parallel On a Pi with under 2GB of RAM the C++ build is limited to fewer parallel
@@ -265,6 +360,7 @@ while [ $# -gt 0 ]; do
--skip-swap) SKIP_SWAP=1 ;; --skip-swap) SKIP_SWAP=1 ;;
--enable-auto-update) AUTO_UPDATE=1 ;; --enable-auto-update) AUTO_UPDATE=1 ;;
--no-auto-update) AUTO_UPDATE=0 ;; --no-auto-update) AUTO_UPDATE=0 ;;
--beta) UPDATE_CHANNEL=beta ;;
--build-jobs) --build-jobs)
shift shift
if [ $# -eq 0 ]; then echo "--build-jobs requires a number"; usage; exit 1; fi if [ $# -eq 0 ]; then echo "--build-jobs requires a number"; usage; exit 1; fi
@@ -291,10 +387,12 @@ else
lm_remove_build_swap() { return 0; } lm_remove_build_swap() { return 0; }
fi fi
# Remove the temporary build swapfile no matter how the script ends. Step 6 # Remove the temporary build swapfile, and take any library patches back out
# tears it down itself; this is the backstop for the error path, since # of the submodule, no matter how the script ends. Step 6 does both itself;
# on_error ends in `exit` and EXIT traps still run. # this is the backstop for the error path (on_error ends in `exit` and EXIT
trap 'lm_remove_build_swap' EXIT # traps still run) and for an interrupted build. _revert_rgb_patches only
# touches patches it applied, so running it twice is harmless.
trap 'lm_remove_build_swap; _revert_rgb_patches' EXIT
# Helpers # Helpers
retry() { retry() {
@@ -873,15 +971,25 @@ if [ -z "$AUTO_UPDATE" ] && [ "$ASSUME_YES" != "1" ] && [ -t 0 ]; then
echo echo
if [[ $REPLY =~ ^[Yy]$ ]]; then AUTO_UPDATE=1; else AUTO_UPDATE=0; fi if [[ $REPLY =~ ^[Yy]$ ]]; then AUTO_UPDATE=1; else AUTO_UPDATE=0; fi
fi fi
if [ "$AUTO_UPDATE" = "1" ] || [ "$AUTO_UPDATE" = "0" ]; then case "$UPDATE_CHANNEL" in
if python3 - "$PROJECT_ROOT_DIR/config/config.json" "$AUTO_UPDATE" <<'PY' stable|beta|"") ;;
*) echo "⚠ LEDMATRIX_CHANNEL=$UPDATE_CHANNEL is not stable or beta; leaving the update channel as it is"
UPDATE_CHANNEL="" ;;
esac
# The update channel, likewise only when asked for (--beta / LEDMATRIX_CHANNEL).
if [ "$AUTO_UPDATE" = "1" ] || [ "$AUTO_UPDATE" = "0" ] || [ -n "$UPDATE_CHANNEL" ]; then
if python3 - "$PROJECT_ROOT_DIR/config/config.json" "$AUTO_UPDATE" "$UPDATE_CHANNEL" <<'PY'
import json, os, sys, tempfile import json, os, sys, tempfile
path, enabled = sys.argv[1], sys.argv[2] == "1" path, enabled = sys.argv[1], sys.argv[2]
channel = sys.argv[3] if len(sys.argv) > 3 else ""
with open(path, encoding="utf-8") as f: with open(path, encoding="utf-8") as f:
config = json.load(f) config = json.load(f)
if not isinstance(config.get("auto_update"), dict): if not isinstance(config.get("auto_update"), dict):
config["auto_update"] = {} config["auto_update"] = {}
config["auto_update"]["enabled"] = enabled if enabled in ("0", "1"):
config["auto_update"]["enabled"] = enabled == "1"
if channel:
config["auto_update"]["channel"] = channel
# Written beside the original and swapped in whole: the display service's # Written beside the original and swapped in whole: the display service's
# config watcher may be running and must never read a half-written file. # config watcher may be running and must never read a half-written file.
original = os.stat(path) original = os.stat(path)
@@ -902,9 +1010,11 @@ except BaseException:
raise raise
PY PY
then then
if [ "$AUTO_UPDATE" = "1" ]; then echo "✓ Weekly automatic updates enabled"; else echo "✓ Weekly automatic updates off"; fi if [ "$AUTO_UPDATE" = "1" ]; then echo "✓ Weekly automatic updates enabled"
elif [ "$AUTO_UPDATE" = "0" ]; then echo "✓ Weekly automatic updates off"; fi
if [ -n "$UPDATE_CHANNEL" ]; then echo "✓ Update channel: $UPDATE_CHANNEL"; fi
else else
echo "⚠ Could not set auto_update in config/config.json; turn it on from the General tab instead" echo "⚠ Could not set auto_update in config/config.json; set it from the General tab instead"
fi fi
fi fi
@@ -1226,9 +1336,11 @@ else
fi fi
BUILD_OUTPUT=$(mktemp) BUILD_OUTPUT=$(mktemp)
BUILD_SUCCESS=false BUILD_SUCCESS=false
_apply_rgb_patches
if run_rgbmatrix_build "$BUILD_JOBS" "$BUILD_OUTPUT"; then if run_rgbmatrix_build "$BUILD_JOBS" "$BUILD_OUTPUT"; then
BUILD_SUCCESS=true BUILD_SUCCESS=true
fi fi
_revert_rgb_patches
cat "$BUILD_OUTPUT" >> "$LOG_FILE" cat "$BUILD_OUTPUT" >> "$LOG_FILE"
if [ "$BUILD_SUCCESS" != true ]; then if [ "$BUILD_SUCCESS" != true ]; then
print_rgbmatrix_build_failure "$BUILD_OUTPUT" print_rgbmatrix_build_failure "$BUILD_OUTPUT"
@@ -1349,15 +1461,16 @@ if ! command -v setcap >/dev/null 2>&1; then
echo "⚠ setcap not found, skipping capability configuration" echo "⚠ setcap not found, skipping capability configuration"
echo " Install libcap2-bin if you need hardware timing capabilities" echo " Install libcap2-bin if you need hardware timing capabilities"
else else
# Find the Python binary and resolve symlinks to get the real binary # The binary the services run (ExecStart=/usr/bin/python3), symlinks
# resolved: python3.11 on Bookworm, python3.13 on Trixie. This used to
# prefer /usr/bin/python3.13 whenever it existed, which would set the
# capability on an interpreter the services never run if python3 pointed
# elsewhere.
PYTHON_BIN="" PYTHON_BIN=""
PYTHON_VER="" PYTHON_VER=""
if [ -f "/usr/bin/python3.13" ]; then if [ -f "/usr/bin/python3" ]; then
PYTHON_BIN=$(readlink -f /usr/bin/python3.13)
PYTHON_VER="3.13"
elif [ -f "/usr/bin/python3" ]; then
PYTHON_BIN=$(readlink -f /usr/bin/python3) PYTHON_BIN=$(readlink -f /usr/bin/python3)
PYTHON_VER=$(python3 --version 2>&1 | grep -oP '(?<=Python )\d+\.\d+' || echo "unknown") PYTHON_VER=$(lm_python_version /usr/bin/python3) || PYTHON_VER="unknown"
fi fi
if [ -n "$PYTHON_BIN" ] && [ -f "$PYTHON_BIN" ]; then if [ -n "$PYTHON_BIN" ] && [ -f "$PYTHON_BIN" ]; then
+5 -4
View File
@@ -6,8 +6,9 @@
files = src files = src
exclude = (^|/)(test|__pycache__)/ exclude = (^|/)(test|__pycache__)/
# Python version # Python version: the oldest the installer supports (Raspberry Pi OS
python_version = 3.10 # Bookworm ships 3.11; Trixie ships 3.13).
python_version = 3.11
# Platform (Linux/Raspberry Pi) # Platform (Linux/Raspberry Pi)
platform = linux platform = linux
@@ -103,8 +104,8 @@ ignore_missing_imports = True
# numpy's own stubs (numpy>=2.3) use Python 3.12 `type` statements, which mypy # numpy's own stubs (numpy>=2.3) use Python 3.12 `type` statements, which mypy
# refuses to parse under python_version = 3.10 -- and 3.10 is the floor this # refuses to parse under python_version = 3.11 -- and 3.11 (Bookworm) is the
# code has to run on, so it stays. Treat numpy as Any instead: skip it, and # floor this code has to run on, so it stays. Treat numpy as Any instead: skip it, and
# follow_imports_for_stubs makes the skip apply to its .pyi files too. # follow_imports_for_stubs makes the skip apply to its .pyi files too.
[mypy-numpy.*] [mypy-numpy.*]
follow_imports = skip follow_imports = skip
@@ -0,0 +1,174 @@
Faster SetImage for rpi-rgb-led-matrix (applied by first_time_install.sh at build
time; the submodule itself stays at its pinned commit).
Copying a frame into the panel buffer was the biggest CPU cost LEDMatrix owns on
large panels: the binding's SetPixelsPillow walked the image column by column
and called SetPixel per pixel, and each SetPixel read-modify-writes one word per
PWM bit plane, 2KB apart, so consecutive pixels were a whole double-row apart and
almost every write missed the cache. This patch:
* FrameCanvas gets its own SetPixelsPillow: row by row, one bulk SetPixels call
per row;
* Framebuffer::SetPixels clips once, looks colours up once per pixel, walks each
row's designators in order and writes the bit planes branch-free;
* the base Canvas.SetPixelsPillow loop (RGBMatrix.SetImage) is row-major.
The bit-plane buffer is byte-identical to the old code's (882 memcmp checks over
noise/gradient/solid/low-value/sparse images, clipped offsets, pwm 7/8/11,
brightness 1/50/90/100, inverse colours, luminance correction off and a pixel
mapper). Measured on a Pi 4 at 512x64: 6.0-6.3 ms -> 1.8 ms per frame through
the Python binding; on hdpi (Pi 4, 4x128x64) frame copy 6.57 -> 2.21 ms and the
display process 139% -> 103% of a core.
LEDMatrix always draws into the canvas that is not on screen and swaps it in
(DisplayManager.update_display), so the write order cannot show as tearing.
Against hzeller/rpi-rgb-led-matrix 1ee4f76.
diff --git a/bindings/python/rgbmatrix/core.pyx b/bindings/python/rgbmatrix/core.pyx
index 230d87f..babc3bb 100644
--- a/bindings/python/rgbmatrix/core.pyx
+++ b/bindings/python/rgbmatrix/core.pyx
@@ -2,6 +2,7 @@
from libcpp cimport bool
from libc.stdint cimport uint8_t, uint32_t, uintptr_t
+from libc.stdlib cimport malloc, free
import cython
cdef extern from "Python.h":
@@ -59,8 +60,9 @@ cdef class Canvas:
buffer = get_pillow_buffer(image_capsule)
- for col in range(max(0, -xstart), min(width, frame_width - xstart)):
- for row in range(max(0, -ystart), min(height, frame_height - ystart)):
+ # Row-major: walks both the image and the bitplane buffer sequentially.
+ for row in range(max(0, -ystart), min(height, frame_height - ystart)):
+ for col in range(max(0, -xstart), min(width, frame_width - xstart)):
pixel = buffer[row][col]
r = (pixel ) & 0xFF
g = (pixel >> 8) & 0xFF
@@ -86,6 +88,41 @@ cdef class FrameCanvas(Canvas):
def SetPixel(self, int x, int y, uint8_t red, uint8_t green, uint8_t blue):
(<cppinc.FrameCanvas*>self._getCanvas()).SetPixel(x, y, red, green, blue)
+ @cython.boundscheck(False)
+ @cython.wraparound(False)
+ def SetPixelsPillow(self, int xstart, int ystart, int width, int height, object image_capsule):
+ # Same result as Canvas.SetPixelsPillow(), but hands each image row
+ # to the C++ bulk FrameCanvas::SetPixels() instead of calling the
+ # virtual SetPixel() once per pixel.
+ cdef cppinc.FrameCanvas* my_canvas = <cppinc.FrameCanvas*>self._getCanvas()
+ cdef int col_start = max(0, -xstart)
+ cdef int col_end = min(width, my_canvas.width() - xstart)
+ cdef int row_start = max(0, -ystart)
+ cdef int row_end = min(height, my_canvas.height() - ystart)
+ cdef int row, col, pixel
+ cdef int *src
+ cdef cppinc.Color *line
+ cdef int **buffer
+
+ if col_end <= col_start or row_end <= row_start:
+ return
+ buffer = get_pillow_buffer(image_capsule)
+ line = <cppinc.Color*>malloc((col_end - col_start) * sizeof(cppinc.Color))
+ if line == NULL:
+ raise MemoryError()
+ try:
+ for row in range(row_start, row_end):
+ src = buffer[row]
+ for col in range(col_start, col_end):
+ pixel = src[col]
+ line[col - col_start].r = pixel & 0xFF
+ line[col - col_start].g = (pixel >> 8) & 0xFF
+ line[col - col_start].b = (pixel >> 16) & 0xFF
+ my_canvas.SetPixels(xstart + col_start, ystart + row,
+ col_end - col_start, 1, line)
+ finally:
+ free(line)
+
property width:
def __get__(self): return (<cppinc.FrameCanvas*>self._getCanvas()).width()
diff --git a/bindings/python/rgbmatrix/cppinc.pxd b/bindings/python/rgbmatrix/cppinc.pxd
index 8bec241..314332d 100644
--- a/bindings/python/rgbmatrix/cppinc.pxd
+++ b/bindings/python/rgbmatrix/cppinc.pxd
@@ -25,6 +25,7 @@ cdef extern from "led-matrix.h" namespace "rgb_matrix":
FrameCanvas *SwapOnVSync(FrameCanvas*, uint8_t)
cdef cppclass FrameCanvas(Canvas):
+ void SetPixels(int, int, int, int, Color*) nogil
bool SetPWMBits(uint8_t)
uint8_t pwmbits()
void SetBrightness(uint8_t)
diff --git a/lib/framebuffer.cc b/lib/framebuffer.cc
index 36d138b..aee62ca 100644
--- a/lib/framebuffer.cc
+++ b/lib/framebuffer.cc
@@ -807,11 +807,60 @@ void Framebuffer::SetPixel(int x, int y, uint8_t r, uint8_t g, uint8_t b) {
}
}
+// Bulk version of SetPixel(); produces exactly the same bitplane content.
+// Faster because it hoists the per-pixel work out of the loop: the color
+// mapping becomes one 256-entry table built per call (each channel maps
+// independently through the same function), the pixel designators of a row
+// are contiguous in the PixelDesignatorMap, and the bit-plane loop is
+// branchless (the color bits are effectively random, so the branches in
+// SetPixel() mispredict a lot).
void Framebuffer::SetPixels(int x, int y, int width, int height, Color *colors) {
- for (int iy = 0; iy < height; ++iy) {
- for (int ix = 0; ix < width; ++ix) {
- SetPixel(x + ix, y + iy, colors->r, colors->g, colors->b);
- ++colors;
+ PixelDesignatorMap *const mapper = *shared_mapper_;
+ const int ix_start = std::max(0, -x);
+ const int ix_end = std::min(width, mapper->width() - x);
+ const int iy_start = std::max(0, -y);
+ const int iy_end = std::min(height, mapper->height() - y);
+ if (ix_start >= ix_end || iy_start >= iy_end) return;
+
+ // Common case (luminance correction, no inversion): use the precomputed
+ // table directly; otherwise build one. Cheap enough to do per call, which
+ // matters for callers that send one row at a time.
+ uint16_t local_map[256];
+ const uint16_t *color_map;
+ if (do_luminance_correct_ && !inverse_color_) {
+ color_map = ColorLookupTable::GetLookup(brightness_).color;
+ } else {
+ for (int c = 0; c < 256; ++c) {
+ uint16_t unused1, unused2;
+ MapColors(c, 0, 0, &local_map[c], &unused1, &unused2);
+ }
+ color_map = local_map;
+ }
+
+ const int min_bit_plane = kBitPlanes - pwm_bits_;
+ gpio_bits_t *const plane_start = bitplane_buffer_ + columns_ * min_bit_plane;
+ for (int iy = iy_start; iy < iy_end; ++iy) {
+ const Color *c = colors + iy * width + ix_start;
+ const PixelDesignator *designator = mapper->get(x + ix_start, y + iy);
+ for (int ix = ix_start; ix < ix_end; ++ix, ++c, ++designator) {
+ const long pos = designator->gpio_word;
+ if (pos < 0) continue; // non-used pixel marker.
+ const uint16_t red = color_map[c->r];
+ const uint16_t green = color_map[c->g];
+ const uint16_t blue = color_map[c->b];
+ const gpio_bits_t r_bits = designator->r_bit;
+ const gpio_bits_t g_bits = designator->g_bit;
+ const gpio_bits_t b_bits = designator->b_bit;
+ const gpio_bits_t designator_mask = designator->mask;
+ gpio_bits_t *bits = plane_start + pos;
+ for (int plane = min_bit_plane; plane < kBitPlanes; ++plane) {
+ const gpio_bits_t color_bits =
+ (r_bits & -(gpio_bits_t)((red >> plane) & 1))
+ | (g_bits & -(gpio_bits_t)((green >> plane) & 1))
+ | (b_bits & -(gpio_bits_t)((blue >> plane) & 1));
+ *bits = (*bits & designator_mask) | color_bits;
+ bits += columns_;
+ }
}
}
}
+1 -1
View File
@@ -1,5 +1,5 @@
# LEDMatrix Core Dependencies # LEDMatrix Core Dependencies
# Compatible with Python 3.10, 3.11, 3.12, and 3.13 # Compatible with Python 3.11, 3.12 and 3.13; CI tests 3.11 and 3.13
# Tested on Raspbian OS 12 (Bookworm) and 13 (Trixie) # Tested on Raspbian OS 12 (Bookworm) and 13 (Trixie)
# Image processing # Image processing
+66 -32
View File
@@ -53,26 +53,35 @@ else
fi fi
echo "" echo ""
# Check OS version # Check OS version. The supported releases come from the same library the
# installer uses, so the two cannot disagree.
echo "2. Checking Operating System Version..." echo "2. Checking Operating System Version..."
echo "---------------------------------------" echo "---------------------------------------"
if [ -f /etc/os-release ]; then OS_LIB="$(cd "$(dirname "$0")" && pwd)/install/lib_os.sh"
. /etc/os-release OS_LIB_LOADED=0
echo "OS: $PRETTY_NAME" OS_RELEASE=""
echo "Version ID: ${VERSION_ID:-unknown}" if [ -f "$OS_LIB" ]; then
# shellcheck source=scripts/install/lib_os.sh
. "$OS_LIB"
OS_LIB_LOADED=1
fi
# first_time_install.sh refuses anything but Raspberry Pi OS / Debian 13 if [ "$OS_LIB_LOADED" = "0" ]; then
# (Trixie), so anything else is an error here too, not a warning. print_error "$OS_LIB is missing - download LEDMatrix again"
if [[ "$ID" == "raspbian" ]] || [[ "$ID" == "debian" ]]; then elif [ -r "$LM_OS_RELEASE_FILE" ]; then
if [ "${VERSION_ID:-0}" = "13" ]; then OS_ID=$(lm_os_field ID)
print_success "Detected Debian 13 Trixie - supported" OS_VERSION_ID=$(lm_os_field VERSION_ID)
elif [ "${VERSION_ID:-0}" = "12" ]; then echo "OS: $(lm_os_field PRETTY_NAME)"
print_error "Debian 12 Bookworm is not supported - the installer requires Raspberry Pi OS Lite (Trixie), Debian 13" echo "Version ID: ${OS_VERSION_ID:-unknown}"
# first_time_install.sh refuses anything else, so this is an error here
# too, not a warning.
if OS_RELEASE=$(lm_os_release); then
print_success "Detected $(lm_release_label "$OS_RELEASE") - supported"
elif [[ "$OS_ID" == "raspbian" ]] || [[ "$OS_ID" == "debian" ]]; then
print_error "Debian/Raspbian ${OS_VERSION_ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12)"
else else
print_error "Debian/Raspbian ${VERSION_ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite (Trixie), Debian 13" print_error "${OS_ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12)"
fi
else
print_error "${ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite (Trixie), Debian 13"
fi fi
else else
print_error "Could not detect OS version" print_error "Could not detect OS version"
@@ -92,7 +101,7 @@ if [ "$KERNEL_MAJOR" -ge "6" ]; then
print_success "Kernel version is compatible (6.x or newer)" print_success "Kernel version is compatible (6.x or newer)"
if [ "$KERNEL_MAJOR" -eq "6" ] && [ "$KERNEL_MINOR" -ge "12" ]; then if [ "$KERNEL_MAJOR" -eq "6" ] && [ "$KERNEL_MINOR" -ge "12" ]; then
print_success "Running latest Trixie kernel (6.12 LTS)" print_success "Running a 6.12 LTS or newer kernel"
fi fi
elif [ "$KERNEL_MAJOR" -eq "5" ] && [ "$KERNEL_MINOR" -ge "10" ]; then elif [ "$KERNEL_MAJOR" -eq "5" ] && [ "$KERNEL_MINOR" -ge "10" ]; then
print_success "Kernel version is compatible (5.10+)" print_success "Kernel version is compatible (5.10+)"
@@ -104,25 +113,34 @@ echo ""
# Check Python version # Check Python version
echo "4. Checking Python Version..." echo "4. Checking Python Version..."
echo "-----------------------------" echo "-----------------------------"
if command -v python3 >/dev/null 2>&1; then if [ "$OS_LIB_LOADED" = "1" ] && command -v python3 >/dev/null 2>&1; then
PYTHON_VERSION=$(python3 -c 'import sys; print(f"{sys.version_info.major}.{sys.version_info.minor}.{sys.version_info.micro}")') PYTHON_VERSION=$(python3 -c 'import sys; print("%d.%d.%d" % sys.version_info[:3])')
PYTHON_MAJOR=$(python3 -c 'import sys; print(sys.version_info.major)') PYTHON_MINOR_VERSION=$(lm_python_version) || PYTHON_MINOR_VERSION=""
PYTHON_MINOR=$(python3 -c 'import sys; print(sys.version_info.minor)') PYTHON_RANGE="3.${LM_PYTHON_MIN_MINOR}-3.${LM_PYTHON_MAX_MINOR}"
echo "Python: $PYTHON_VERSION" echo "Python: $PYTHON_VERSION"
if [ "$PYTHON_MAJOR" -eq "3" ]; then case "$(lm_python_check "$PYTHON_MINOR_VERSION")" in
if [ "$PYTHON_MINOR" -ge "10" ] && [ "$PYTHON_MINOR" -le "13" ]; then ok)
print_success "Python version is supported (3.10-3.13)" print_success "Python version is supported ($PYTHON_RANGE)"
elif [ "$PYTHON_MINOR" -ge "14" ]; then ;;
print_warning "Python 3.${PYTHON_MINOR} is very new - some packages may not be compatible yet" too-old)
else # The rgbmatrix bindings declare requires-python >=3.11, so the
# Pillow 12 and the pinned test tools need 3.10+, so this won't install. # display cannot be built on anything older.
print_error "Python 3.${PYTHON_MINOR} is too old - Python 3.10+ is required" print_error "Python $PYTHON_MINOR_VERSION is too old - Python 3.${LM_PYTHON_MIN_MINOR}+ is required"
fi ;;
else too-new)
print_error "Python 2.x detected - Python 3.10+ is required" print_warning "Python $PYTHON_MINOR_VERSION is newer than LEDMatrix has been tested with ($PYTHON_RANGE)"
;;
*)
print_warning "Could not read the Python version"
;;
esac
if [ -n "$OS_RELEASE" ] && [ "$PYTHON_MINOR_VERSION" != "$(lm_release_python "$OS_RELEASE")" ]; then
print_warning "$(lm_release_label "$OS_RELEASE") ships Python $(lm_release_python "$OS_RELEASE"), but python3 runs $PYTHON_MINOR_VERSION"
fi fi
elif command -v python3 >/dev/null 2>&1; then
print_warning "Cannot check the Python version without $OS_LIB"
else else
print_error "Python 3 not found - installation required" print_error "Python 3 not found - installation required"
fi fi
@@ -268,6 +286,22 @@ if command -v ping >/dev/null 2>&1; then
else else
print_warning "Ping command not available - cannot verify network" print_warning "Ping command not available - cannot verify network"
fi fi
# WiFi setup from the web page and the LEDMatrix-Setup hotspot drive
# NetworkManager, the default on both Bookworm and Trixie.
if [ "$OS_LIB_LOADED" = "1" ]; then
case "$(lm_network_stack)" in
networkmanager)
print_success "NetworkManager manages the network (needed for WiFi setup)"
;;
dhcpcd)
print_warning "dhcpcd manages the network - WiFi setup from the web page and the setup hotspot need NetworkManager (sudo raspi-config -> Advanced Options -> Network Config)"
;;
*)
print_warning "Could not tell which service manages the network - WiFi setup from the web page needs NetworkManager"
;;
esac
fi
echo "" echo ""
# Print summary # Print summary
+13 -3
View File
@@ -5,11 +5,18 @@ This directory contains scripts for installing and configuring the LEDMatrix sys
## Scripts ## Scripts
- **`one-shot-install.sh`** - Single-command installer; clones the - **`one-shot-install.sh`** - Single-command installer; clones the
repo, checks prerequisites, then runs `first_time_install.sh`. repo, checks out the newest release (or `main` with
Invoked via `curl ... | bash` from the project root README. `LEDMATRIX_CHANNEL=beta`), checks prerequisites, then runs
`first_time_install.sh`. Invoked via `curl ... | bash` from the project
root README. Re-running it never moves a checkout to an older version.
- **`install_service.sh`** - Installs, enables and starts the display - **`install_service.sh`** - Installs, enables and starts the display
service (`ledmatrix.service`), the web interface service service (`ledmatrix.service`), the web interface service
(`ledmatrix-web.service`) and the update-verify units (systemd) (`ledmatrix-web.service`) and the update-verify units (systemd), and
installs `/usr/local/sbin/ledmatrix-refresh-units`
- **`ledmatrix_refresh_units.py`** - Not run from here: `install_service.sh`
installs a root-owned copy as `/usr/local/sbin/ledmatrix-refresh-units`,
which updates run through sudo to install changed units (and the
automatic update's rollback, with `--restore`, to put them back)
- **`install_web_service.sh`** - Installs only the web interface service - **`install_web_service.sh`** - Installs only the web interface service
and the update-verify units (systemd) and the update-verify units (systemd)
- **`install_wifi_monitor.sh`** - Installs the WiFi monitor daemon service - **`install_wifi_monitor.sh`** - Installs the WiFi monitor daemon service
@@ -34,6 +41,9 @@ Libraries (sourced, not run):
script that renders a unit from `systemd/*.service` script that renders a unit from `systemd/*.service`
- **`lib_lowmem.sh`** - Build-job sizing and temporary swap for the C++ - **`lib_lowmem.sh`** - Build-job sizing and temporary swap for the C++
build on low-memory Pis (`first_time_install.sh` Step 6) build on low-memory Pis (`first_time_install.sh` Step 6)
- **`lib_os.sh`** - Which releases (Bookworm, Trixie) and Python versions
(3.11-3.13) the installer accepts, and which service runs the network;
shared by `first_time_install.sh` and `scripts/check_system_compatibility.sh`
## Usage ## Usage
+2
View File
@@ -138,6 +138,8 @@ echo "- View system logs via journalctl"
echo "- Reboot and shutdown the system" echo "- Reboot and shutdown the system"
echo "- Remove plugin directories (for update/uninstall when root-owned files block deletion)" echo "- Remove plugin directories (for update/uninstall when root-owned files block deletion)"
echo "- Install plugin/base requirements.txt as root (so ledmatrix.service can see them)" echo "- Install plugin/base requirements.txt as root (so ledmatrix.service can see them)"
echo "- Install the LEDMatrix systemd units an update changed, and restore them on rollback"
echo " (/usr/local/sbin/ledmatrix-refresh-units, installed by install_service.sh)"
echo "" echo ""
# Ask for confirmation # Ask for confirmation
+24
View File
@@ -143,6 +143,30 @@ for VERIFY_UNIT in ledmatrix-update-verify.service ledmatrix-update-verify.path;
fi fi
done done
# The helper updates run (through sudo, see lib_sudoers.sh) to install these
# same units when a new version changes their templates, and to put the old
# ones back if the automatic update rolls back. Root-owned and outside the
# checkout, so the web user who owns the checkout cannot change what sudo runs.
# Not fatal: without it, updates leave the units for the next reinstall.
REFRESH_UNITS_SRC="$PROJECT_ROOT_DIR/scripts/install/ledmatrix_refresh_units.py"
REFRESH_UNITS_DEST=/usr/local/sbin/ledmatrix-refresh-units
if [ -f "$REFRESH_UNITS_SRC" ]; then
if sudo install -D -o root -g root -m 0755 "$REFRESH_UNITS_SRC" "$REFRESH_UNITS_DEST"; then
echo "Installed $REFRESH_UNITS_DEST (lets updates refresh these units)"
else
echo "WARNING: could not install $REFRESH_UNITS_DEST; updates will not refresh the systemd units" >&2
fi
fi
# The units above are copied from mktemp files, which are 0600. 0644 is what
# first_time_install.sh (Step 8.1) sets, and lets the web interface compare
# them with the templates after an update without root.
for INSTALLED_UNIT in ledmatrix.service ledmatrix-web.service \
ledmatrix-update-verify.service ledmatrix-update-verify.path; do
if [ -f "/etc/systemd/system/$INSTALLED_UNIT" ]; then
sudo chmod 644 "/etc/systemd/system/$INSTALLED_UNIT" || true
fi
done
echo "Reloading systemd daemon for web service..." echo "Reloading systemd daemon for web service..."
sudo systemctl daemon-reload sudo systemctl daemon-reload
+451
View File
@@ -0,0 +1,451 @@
#!/usr/bin/python3 -I
"""Refresh the installed LEDMatrix systemd units from the checkout's templates.
Installed by scripts/install/install_service.sh as a root-owned copy,
/usr/local/sbin/ledmatrix-refresh-units, and granted to the web interface's
user by /etc/sudoers.d/ledmatrix_web (scripts/install/lib_sudoers.sh) with
exactly two command lines:
ledmatrix-refresh-units (no arguments)
ledmatrix-refresh-units --restore
An update (Update Code, or the weekly automatic update) pulls new unit
templates into systemd/, but the units systemd runs are the copies in
/etc/systemd/system, which only the installer used to write. So a setting
added to a template -- the render-loop watchdog, a memory limit -- never
reached a device that was already installed. After an update the web
interface runs this, and the next restart picks the new units up.
* **No arguments:** render each installed unit from systemd/<unit> exactly as
install_service.sh does (__PROJECT_ROOT_DIR__ and __USER__ replaced
literally), and install the ones whose content differs (comments and blank
lines aside, as src/startup_validator.py compares them), then
``systemctl daemon-reload``. The units replaced are saved first, so the
automatic update's rollback can put them back.
* ``--restore``: put back the units the last refresh replaced, and
daemon-reload. Nothing saved means nothing to do.
* ``--check``: print the units that would change, one per line. Needs no
root and changes nothing.
What it trusts, and why. It takes no other input: the project directory and
the web interface's user come from the installed, root-owned
ledmatrix.service and ledmatrix-web.service, not from the caller, and sudo
strips the caller's environment (``-I`` ignores the PYTHON* variables too).
It only replaces units that are already installed, only the four
install_service.sh installs, and only with a rendering that keeps each unit's
User= (root for the display, the web user for the others) and
WorkingDirectory=. The templates are files the web user can edit -- but so is
run.py, which ledmatrix.service already runs as root, so a template grants
nothing that user did not have; the checks keep a damaged or hostile template
from changing who a unit runs as, and keep this from reading anything but a
regular file under the checkout's systemd/ folder.
Standard library only, and no imports from the checkout: the installed copy
must not run code the web user can change.
"""
import json
import os
import re
import stat
import subprocess # nosec B404 - fixed argv, no shell # nosemgrep
import sys
import tempfile
SYSTEMD_DIR = '/etc/systemd/system'
#: Root-only: the units the last refresh replaced, for --restore.
BACKUP_DIR = '/var/lib/ledmatrix/unit-backup'
MANIFEST = 'manifest.json'
INSTALLED_PATH = '/usr/local/sbin/ledmatrix-refresh-units'
DISPLAY_UNIT = 'ledmatrix.service'
WEB_UNIT = 'ledmatrix-web.service'
VERIFY_SERVICE = 'ledmatrix-update-verify.service'
VERIFY_PATH = 'ledmatrix-update-verify.path'
#: What install_service.sh installs, in its order. Nothing else is touched.
UNITS = (DISPLAY_UNIT, WEB_UNIT, VERIFY_SERVICE, VERIFY_PATH)
MAX_TEMPLATE_BYTES = 64 * 1024
_USER_RE = re.compile(r'^[a-z_][a-z0-9_-]{0,31}$')
#: systemd expands % specifiers, and a quote, backslash or line break would
#: be reinterpreted in a unit file (src/auto_update_setup.py refuses the same).
#: (On Windows, where the tests also run, a backslash is the path separator.)
_UNSAFE_PATH_CHARS = set('%"') | ({'\\'} if os.sep == '/' else set())
EXIT_OK = 0
EXIT_FAILED = 1
EXIT_USAGE = 2
class RefreshError(Exception):
"""Why the units were left alone, in words for the web interface's log."""
class UnitsUnreadable(RefreshError):
"""An installed unit is not readable by this (unprivileged) user.
install_service.sh used to leave units mode 0600 (first_time_install.sh's
Step 8.1 makes them 0644), so ``--check`` as the web user cannot always
tell; the root helper itself can.
"""
def directive_values(text, key):
"""Every value of ``key=`` in a unit's text, in order (systemd allows spaces around ``=``)."""
return [m.group(1).strip() for m in re.finditer(rf'^[ \t]*{key}[ \t]*=(.*)$', text or '', re.M)]
def layout_problem(text, section, keys):
"""What would make ``directive_values`` misread the unit as systemd reads it, or None.
A ``User=`` inside a backslash-continued line is part of the line before,
and one under [Unit] is ignored, so either could pass a check that systemd
then does not apply. Neither appears in the shipped templates.
"""
current = None
for raw in (text or '').splitlines():
line = raw.strip()
if not line or line.startswith(('#', ';')):
continue
if line.endswith('\\'):
return 'continues a line with a backslash'
if line.startswith('[') and line.endswith(']'):
current = line[1:-1]
continue
key = line.split('=', 1)[0].strip()
if key in keys and current != section:
return f'sets {key}= outside [{section}]'
return None
def unit_body(text):
"""A unit's meaningful lines in order: no comments, no blank lines.
The same comparison src/startup_validator.py uses for its drift warning,
so what this refreshes is exactly what that warns about.
"""
lines = []
for line in (text or '').splitlines():
line = line.strip()
if line and not line.startswith('#'):
lines.append(line)
return '\n'.join(lines)
def render(template, project_root, user):
"""install_service.sh's ``sed "s|__PROJECT_ROOT_DIR__|...|g; s|__USER__|...|g"``."""
return template.replace('__PROJECT_ROOT_DIR__', project_root).replace('__USER__', user)
def _read_regular(path, limit=MAX_TEMPLATE_BYTES, dir_fd=None):
"""A regular file's text, never through a symlink, a FIFO or a device."""
flags = os.O_RDONLY | getattr(os, 'O_NOFOLLOW', 0) | getattr(os, 'O_NONBLOCK', 0)
kwargs = {'dir_fd': dir_fd} if dir_fd is not None else {}
fd = os.open(path, flags, **kwargs)
try:
info = os.fstat(fd)
if not stat.S_ISREG(info.st_mode):
raise RefreshError(f'{path} is not a regular file')
if info.st_size > limit:
raise RefreshError(f'{path} is larger than {limit} bytes')
data = b''
while True:
chunk = os.read(fd, limit + 1 - len(data))
if not chunk:
break
data += chunk
if len(data) > limit:
raise RefreshError(f'{path} is larger than {limit} bytes')
finally:
os.close(fd)
if b'\0' in data:
raise RefreshError(f'{path} is not a text file')
try:
return data.decode('utf-8')
except UnicodeDecodeError as e:
raise RefreshError(f'{path} is not UTF-8') from e
def _read_installed(systemd_dir, name):
path = os.path.join(systemd_dir, name)
try:
with open(path, 'r', encoding='utf-8') as f:
return f.read()
except FileNotFoundError:
return None
except PermissionError as e:
raise UnitsUnreadable(f'cannot read the installed {name}: {e}') from e
except (OSError, UnicodeDecodeError) as e:
raise RefreshError(f'cannot read the installed {name}: {e}') from e
def _lookup_user(user):
try:
import pwd
except ImportError: # not a POSIX host (the tests on Windows)
return True
try:
pwd.getpwnam(user)
return True
except KeyError:
return False
class Refresher:
def __init__(self, systemd_dir=SYSTEMD_DIR, backup_dir=BACKUP_DIR, run=subprocess.run,
is_root=None, user_exists=_lookup_user, log=None):
self.systemd_dir = systemd_dir
self.backup_dir = backup_dir
self.run = run
self.is_root = is_root or (lambda: hasattr(os, 'geteuid') and os.geteuid() == 0)
self.user_exists = user_exists
self.log = log or (lambda msg: print(msg, flush=True))
# -- what the installed units say -------------------------------------
def context(self, installed):
"""(project root, web user) from the installed, root-owned units."""
display = installed.get(DISPLAY_UNIT)
if display is None:
raise RefreshError(f'{DISPLAY_UNIT} is not installed; run scripts/install/install_service.sh')
roots = directive_values(display, 'WorkingDirectory')
if len(roots) != 1:
raise RefreshError(f'the installed {DISPLAY_UNIT} does not name one WorkingDirectory')
root = roots[0]
if (not os.path.isabs(root) or any(ch in _UNSAFE_PATH_CHARS or ord(ch) < 32 for ch in root)
or os.path.normpath(root) != root):
raise RefreshError(f'the installed {DISPLAY_UNIT} runs from {root!r}, which cannot be used')
if not os.path.isdir(root):
raise RefreshError(f'{root} (the installed {DISPLAY_UNIT} WorkingDirectory) does not exist')
user = None
web = installed.get(WEB_UNIT)
if web is not None:
users = directive_values(web, 'User')
user = users[0] if len(users) == 1 else ('root' if not users else None)
if user is None or not _USER_RE.match(user) or not self.user_exists(user):
raise RefreshError(f'the installed {WEB_UNIT} runs as an account that cannot be used')
if directive_values(web, 'WorkingDirectory') != [root]:
raise RefreshError(f'the installed {WEB_UNIT} and {DISPLAY_UNIT} run from different folders')
return root, user
@staticmethod
def expected_user(name, web_user):
return 'root' if name == DISPLAY_UNIT else web_user
def _template(self, root, name):
"""systemd/<name> under the checkout, as a regular file, never via a symlink."""
dir_flags = os.O_RDONLY | getattr(os, 'O_DIRECTORY', 0) | getattr(os, 'O_NOFOLLOW', 0)
if os.open in getattr(os, 'supports_dir_fd', set()):
try:
dfd = os.open(os.path.join(root, 'systemd'), dir_flags)
except OSError as e:
raise RefreshError(f'cannot open {root}/systemd: {e}') from e
try:
return _read_regular(name, dir_fd=dfd)
except FileNotFoundError:
return None
except OSError as e:
raise RefreshError(f'cannot read systemd/{name}: {e}') from e
finally:
os.close(dfd)
path = os.path.join(root, 'systemd', name)
if os.path.islink(os.path.join(root, 'systemd')):
raise RefreshError(f'{root}/systemd is a symlink')
try:
return _read_regular(path)
except FileNotFoundError:
return None
except OSError as e:
raise RefreshError(f'cannot read systemd/{name}: {e}') from e
def _validate(self, name, rendered, root, user):
problem = layout_problem(rendered, 'Service', ('User', 'WorkingDirectory'))
if problem:
raise RefreshError(f'systemd/{name} {problem}; refusing to install it')
if directive_values(rendered, 'User') != [user]:
raise RefreshError(f'systemd/{name} would not run as {user}; refusing to install it')
if directive_values(rendered, 'WorkingDirectory') != [root]:
raise RefreshError(f'systemd/{name} would not run from {root}; refusing to install it')
def plan(self):
"""{unit: (installed text, new text)} for every installed unit that would change.
Raises RefreshError, and so changes nothing, if any unit cannot be
rendered safely: four units refreshed as a set or not at all.
"""
installed = {name: _read_installed(self.systemd_dir, name) for name in UNITS}
root, web_user = self.context(installed)
changes = {}
for name in UNITS:
current = installed[name]
if current is None:
continue # never installed here: installing is the installer's job
user = self.expected_user(name, web_user)
if user is None:
continue # the web unit is not installed, so neither is its user
template = self._template(root, name)
if template is None:
continue # a version without this unit leaves the installed one alone
rendered = render(template, root, user)
# A path unit runs nothing itself; what matters is what it starts.
if name.endswith('.service'):
self._validate(name, rendered, root, user)
else:
self._validate_path(name, rendered)
if unit_body(rendered) != unit_body(current):
changes[name] = (current, rendered)
return changes
def _validate_path(self, name, rendered):
problem = layout_problem(rendered, 'Path', ('Unit',))
if problem:
raise RefreshError(f'systemd/{name} {problem}; refusing to install it')
if directive_values(rendered, 'Unit') != [VERIFY_SERVICE]:
raise RefreshError(f'systemd/{name} does not start {VERIFY_SERVICE}; refusing to install it')
if directive_values(rendered, 'User'):
raise RefreshError(f'systemd/{name} sets User=; refusing to install it')
# -- writing ------------------------------------------------------------
def _write_unit(self, name, text):
fd, tmp = tempfile.mkstemp(dir=self.systemd_dir, prefix=f'.{name}.')
try:
with os.fdopen(fd, 'w', encoding='utf-8', newline='\n') as f:
f.write(text)
os.chmod(tmp, 0o644)
os.replace(tmp, os.path.join(self.systemd_dir, name))
except BaseException:
try:
os.unlink(tmp)
except OSError:
pass
raise
def _backup_dir(self):
"""The backup folder, created root-only; refused if it is not a plain folder."""
os.makedirs(os.path.dirname(self.backup_dir), mode=0o755, exist_ok=True)
try:
os.mkdir(self.backup_dir, 0o700)
except FileExistsError:
pass
info = os.lstat(self.backup_dir)
if not stat.S_ISDIR(info.st_mode):
raise RefreshError(f'{self.backup_dir} is not a folder')
if hasattr(os, 'geteuid') and info.st_uid != os.geteuid():
raise RefreshError(f'{self.backup_dir} is not owned by root')
return self.backup_dir
def _clear_backup(self, folder):
for entry in os.listdir(folder):
path = os.path.join(folder, entry)
if os.path.isfile(path) or os.path.islink(path):
os.unlink(path)
def _systemctl(self, *args):
result = self.run(['systemctl', *args], capture_output=True, text=True, timeout=60)
if result.returncode != 0:
raise RefreshError(f'"systemctl {" ".join(args)}" failed: '
f'{(result.stderr or result.stdout or "").strip()}')
def _restart_path_unit_if_active(self, names):
"""A rewritten path unit watches the old path until it is restarted."""
if VERIFY_PATH not in names:
return
state = self.run(['systemctl', 'is-active', VERIFY_PATH], capture_output=True, text=True, timeout=30)
if (state.stdout or '').strip() == 'active':
self._systemctl('restart', VERIFY_PATH)
def refresh(self):
if not self.is_root():
raise RefreshError('must run as root (sudo)')
changes = self.plan()
folder = self._backup_dir()
# Always reset: the backup belongs to this refresh, so a --restore
# after an update that changed nothing restores nothing.
self._clear_backup(folder)
if not changes:
self.log('units: up to date')
return []
for name, (current, _) in changes.items():
with open(os.path.join(folder, name), 'w', encoding='utf-8', newline='\n') as f:
f.write(current)
with open(os.path.join(folder, MANIFEST), 'w', encoding='utf-8') as f:
json.dump({'units': sorted(changes)}, f)
try:
for name, (_, rendered) in changes.items():
self._write_unit(name, rendered)
self._systemctl('daemon-reload')
except BaseException:
# A failed refresh is reported as a failure, so the update records
# no units_refreshed and a rollback would not --restore. Put the
# replaced units back now, rather than leave a half-written set
# under the old code.
self._undo(changes, folder)
raise
self._restart_path_unit_if_active(changes)
self.log('units refreshed: ' + ' '.join(sorted(changes)))
return sorted(changes)
def _undo(self, changes, folder):
"""Best effort: reinstall the units a failed refresh replaced."""
undone = True
for name, (current, _) in changes.items():
try:
self._write_unit(name, current)
except OSError as e:
undone = False
self.log(f'units: could not put back {name}: {e}')
try:
self.run(['systemctl', 'daemon-reload'], capture_output=True, text=True, timeout=60)
except (OSError, subprocess.SubprocessError) as e:
self.log(f'units: daemon-reload after putting units back failed: {e}')
if undone:
# Nothing is left to restore; keep the backup only if a unit could
# not be put back, so a manual --restore still can.
self._clear_backup(folder)
def restore(self):
if not self.is_root():
raise RefreshError('must run as root (sudo)')
folder = self._backup_dir()
try:
manifest = json.loads(_read_regular(os.path.join(folder, MANIFEST)))
except FileNotFoundError:
self.log('units: nothing to restore')
return []
names = [n for n in (manifest or {}).get('units', []) if n in UNITS]
for name in names:
self._write_unit(name, _read_regular(os.path.join(folder, name)))
self._systemctl('daemon-reload')
self._restart_path_unit_if_active(names)
self._clear_backup(folder)
self.log('units restored: ' + ' '.join(names))
return names
def main(argv, refresher=None):
args = argv[1:]
if args not in ([], ['--restore'], ['--check']):
print('usage: ledmatrix-refresh-units [--restore | --check]', file=sys.stderr)
return EXIT_USAGE
refresher = refresher or Refresher()
try:
if args == ['--check']:
for name in sorted(refresher.plan()):
print(name)
elif args == ['--restore']:
refresher.restore()
else:
refresher.refresh()
except (RefreshError, OSError, subprocess.SubprocessError, ValueError) as e:
print(f'ledmatrix-refresh-units: {e}', file=sys.stderr)
return EXIT_FAILED
return EXIT_OK
if __name__ == '__main__':
# Only as the installed program: sudo already sets a secure PATH, and
# this pins the one systemctl comes from. (Not in main(), which the
# tests call in-process.)
os.environ['PATH'] = '/usr/sbin:/usr/bin:/sbin:/bin'
sys.exit(main(sys.argv))
+138
View File
@@ -0,0 +1,138 @@
#!/bin/bash
# Which operating systems and Python versions LEDMatrix installs on.
#
# Sourced by first_time_install.sh and scripts/check_system_compatibility.sh,
# so the installer and the compatibility checker cannot disagree about what
# is supported. Pure functions: nothing here installs, changes or exits --
# the callers decide what to do with the answers.
#
# Supported (Lite, no desktop):
# Raspberry Pi OS / Debian 12 "Bookworm" -- Python 3.11
# Raspberry Pi OS / Debian 13 "Trixie" -- Python 3.13
#
# Everything the installer asks apt for (python3-pip, python3-venv,
# python-dev-is-python3, python3-pil, python3-pil.imagetk, build-essential,
# python3-setuptools, python3-wheel, cmake, ninja-build, git, curl, wget,
# unzip, and hostapd, dnsmasq, network-manager for WiFi setup) has the same
# name on both releases. Both ship a pip (23.0.1 and 25.1.1) that is PEP 668
# "externally managed" and accepts --break-system-packages, and a cmake (3.25
# and 3.31) new enough for the rgbmatrix build (3.22). So no step needs a
# per-release branch today; if one ever does, the release name comes from
# lm_os_release below.
# Test hook: the os-release file to read.
LM_OS_RELEASE_FILE="${LM_OS_RELEASE_FILE:-/etc/os-release}"
# Oldest and newest python3 minor versions the installer accepts. 3.11 is
# Bookworm's, and also the floor of the rgbmatrix bindings (requires-python
# >=3.11 in rpi-rgb-led-matrix-master/pyproject.toml); 3.13 is Trixie's.
LM_PYTHON_MIN_MINOR=11
LM_PYTHON_MAX_MINOR=13
# lm_os_field KEY -- one value from os-release with its quotes removed; empty
# when the key or the file is missing. Parsed rather than sourced so that
# os-release's ID, VERSION and friends do not land in the caller's variables.
lm_os_field() {
[ -r "$LM_OS_RELEASE_FILE" ] || return 0
sed -n "/^$1=/{s/^$1=//;s/^[\"']//;s/[\"']\$//;p;q;}" "$LM_OS_RELEASE_FILE"
}
# lm_os_release -- print "bookworm" or "trixie" and succeed on a supported
# release; print nothing and fail on anything else. VERSION_ID decides; the
# codename is used only when VERSION_ID is missing.
lm_os_release() {
local id version
id=$(lm_os_field ID)
version=$(lm_os_field VERSION_ID)
[ -n "$version" ] || version=$(lm_os_field VERSION_CODENAME)
case "$id" in
raspbian|debian) ;;
*) return 1 ;;
esac
case "$version" in
12|bookworm) echo bookworm ;;
13|trixie) echo trixie ;;
*) return 1 ;;
esac
}
# lm_release_label RELEASE -- how to name a release to a person.
lm_release_label() {
case "$1" in
bookworm) echo "Debian 12 (Bookworm)" ;;
trixie) echo "Debian 13 (Trixie)" ;;
*) echo "$1" ;;
esac
}
# lm_release_python RELEASE -- the python3 version a release ships, e.g. 3.11.
lm_release_python() {
case "$1" in
bookworm) echo 3.11 ;;
trixie) echo 3.13 ;;
*) return 1 ;;
esac
}
# lm_python_version [PYTHON] -- "3.11" and so on for python3 (or PYTHON);
# prints nothing and fails when it cannot be run.
lm_python_version() {
"${1:-python3}" -c 'import sys; print("%d.%d" % sys.version_info[:2])' 2>/dev/null
}
# lm_python_check VERSION -- print "ok", "too-old", "too-new" or "unknown"
# for a version such as 3.11. Always succeeds, so it is safe under set -e.
lm_python_check() {
local major minor
major=${1%%.*}
minor=${1#*.}
minor=${minor%%.*}
case "$major:$minor" in
*[!0-9:]*|:*|*:) echo unknown; return 0 ;;
esac
if [ "$major" -lt 3 ] || { [ "$major" -eq 3 ] && [ "$minor" -lt "$LM_PYTHON_MIN_MINOR" ]; }; then
echo too-old
elif [ "$major" -gt 3 ] || [ "$minor" -gt "$LM_PYTHON_MAX_MINOR" ]; then
echo too-new
else
echo ok
fi
}
# lm_network_stack -- which service runs the network: "networkmanager",
# "dhcpcd" or "unknown". Raspberry Pi OS uses NetworkManager on both Bookworm
# and Trixie; dhcpcd appears when someone switched back to it in raspi-config.
lm_network_stack() {
if systemctl is-active --quiet NetworkManager 2>/dev/null; then
echo networkmanager
elif systemctl is-active --quiet dhcpcd 2>/dev/null; then
echo dhcpcd
else
echo unknown
fi
}
# lm_print_dhcpcd_advice -- the explanation for a Pi running dhcpcd. WiFi
# setup from the web page and the LEDMatrix-Setup hotspot both drive
# NetworkManager (nmcli). The installer does not switch the network stack
# itself: doing that over SSH can cut the connection it is running on.
lm_print_dhcpcd_advice() {
echo "⚠ This Pi manages its network with dhcpcd, not NetworkManager."
echo " LEDMatrix installs and the display works, but choosing a WiFi network"
echo " from the web page and the LEDMatrix-Setup hotspot both need NetworkManager."
echo " To switch (with a keyboard and screen attached, or over Ethernet):"
echo " sudo raspi-config -> Advanced Options -> Network Config -> NetworkManager"
echo " then reboot."
}
# lm_print_supported_os_help -- what to do on an unsupported system.
lm_print_supported_os_help() {
echo "LEDMatrix needs Raspberry Pi OS Lite: Trixie (Debian 13) or Bookworm (Debian 12)."
echo ""
echo "To install Raspberry Pi OS Lite:"
echo " 1. Download Raspberry Pi Imager from: https://www.raspberrypi.com/software/"
echo " 2. Choose 'Raspberry Pi OS Lite (64-bit)'. Trixie is the current version and"
echo " is recommended; Bookworm (listed as Legacy) also works"
echo " 3. Flash it to the SD card"
echo " 4. Boot the Pi and run this script again"
}
+9
View File
@@ -10,6 +10,11 @@
# #
# Add or remove a grant here and nowhere else. # Add or remove a grant here and nowhere else.
# Root-owned copy of scripts/install/ledmatrix_refresh_units.py, installed by
# install_service.sh. Outside the checkout on purpose: the web user owns the
# checkout, so a granted file inside it could be rewritten and run as root.
LEDMATRIX_REFRESH_UNITS_PATH=/usr/local/sbin/ledmatrix-refresh-units
# web_sudoers_rules WEB_USER PROJECT_ROOT SYSTEMCTL_PATH BASH_PATH REBOOT_PATH POWEROFF_PATH JOURNALCTL_PATH # web_sudoers_rules WEB_USER PROJECT_ROOT SYSTEMCTL_PATH BASH_PATH REBOOT_PATH POWEROFF_PATH JOURNALCTL_PATH
# #
# Print the ledmatrix_web sudoers rules to stdout. # Print the ledmatrix_web sudoers rules to stdout.
@@ -58,6 +63,10 @@ $WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/scripts/fix_perms/safe_pl
# Install a requirements.txt as root via vetted helper, so packages are visible # Install a requirements.txt as root via vetted helper, so packages are visible
# to root-run ledmatrix.service (not just the web interface's own user). # to root-run ledmatrix.service (not just the web interface's own user).
$WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/scripts/fix_perms/safe_pip_install.sh * $WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/scripts/fix_perms/safe_pip_install.sh *
# After an update, install the new systemd units (no arguments: "" allows none)
# and, on the automatic update's rollback, put the previous ones back.
$WEB_USER ALL=(ALL) NOPASSWD: $LEDMATRIX_REFRESH_UNITS_PATH ""
$WEB_USER ALL=(ALL) NOPASSWD: $LEDMATRIX_REFRESH_UNITS_PATH --restore
EOF EOF
if [ -n "$JOURNALCTL_PATH" ]; then if [ -n "$JOURNALCTL_PATH" ]; then
cat << EOF cat << EOF
+119 -1
View File
@@ -3,6 +3,10 @@
# LED Matrix One-Shot Installation Script # LED Matrix One-Shot Installation Script
# This script provides a single-command installation experience # This script provides a single-command installation experience
# Usage: curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | bash # Usage: curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | bash
#
# A new install runs the newest release (the stable update channel). For the
# newest code from main instead (the beta channel), set LEDMATRIX_CHANNEL=beta:
# curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | LEDMATRIX_CHANNEL=beta bash
set -Eeuo pipefail set -Eeuo pipefail
@@ -205,6 +209,114 @@ check_sudo() {
print_success "Sudo access confirmed" print_success "Sudo access confirmed"
} }
# --- release checkout helpers ------------------------------------------------
# Which version an install runs. The rules are web_interface/update_channel.py's,
# so the installer and the web interface's updates agree:
# stable (default) the newest vX.Y.Z tag by semantic version; pre-releases
# (v3.8.0-rc1), leading zeros and other tags are ignored
# beta main, the newest code
# Never backwards: an existing checkout moves to a release only when that
# release contains its current commit (git merge-base --is-ancestor).
# Never fatal: whatever goes wrong, the install carries on with the checkout
# as it is.
# Print "stable" or "beta": LEDMATRIX_CHANNEL when it is set, else the
# existing install's auto_update.channel (CONFIG_FILE), else stable.
_lm_channel() {
local config_file="${1:-}" value
value=$(printf '%s' "${LEDMATRIX_CHANNEL:-}" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')
case "$value" in
stable|beta) printf '%s\n' "$value"; return 0 ;;
"") ;;
*) print_warning "LEDMATRIX_CHANNEL=${LEDMATRIX_CHANNEL} is not stable or beta; using stable" >&2
printf 'stable\n'; return 0 ;;
esac
if [ -n "$config_file" ] && [ -f "$config_file" ] && command -v python3 >/dev/null 2>&1; then
value=$(python3 - "$config_file" 2>/dev/null <<'PY' || true
import json, sys
try:
with open(sys.argv[1], encoding="utf-8") as f:
section = json.load(f).get("auto_update")
value = section.get("channel") if isinstance(section, dict) else None
print(value.strip().lower() if isinstance(value, str) else "")
except Exception:
print("")
PY
)
if [ "$value" = "beta" ]; then
printf 'beta\n'
return 0
fi
fi
printf 'stable\n'
}
# Print the newest release tag of the repository in the current directory,
# or nothing when it has none.
_lm_newest_release_tag() {
git tag --list 'v*' 2>/dev/null \
| grep -E '^v(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)$' \
| sort -t. -k1.2,1n -k2,2n -k3,3n \
| tail -n 1 || true
}
# A fresh clone (on main): move to the newest release unless beta was asked for.
_lm_checkout_release_after_clone() {
local channel tag
channel=$(_lm_channel "")
if [ "$channel" = "beta" ]; then
print_success "Beta channel: installing the newest code from main"
return 0
fi
tag=$(_lm_newest_release_tag)
if [ -z "$tag" ]; then
print_warning "No release found; installing the newest code from main"
return 0
fi
if git -c advice.detachedHead=false checkout --quiet --detach "${tag}^{commit}"; then
print_success "Installing release $tag (stable channel)"
else
print_warning "Could not check out release $tag; installing the newest code from main"
fi
return 0
}
# An existing checkout: move it forward along its channel, never backwards.
# Returns 1 when it should be updated the way it always was (a fast-forward
# pull of its branch): beta, or stable on a branch newer than every release.
_lm_update_existing_checkout() {
local channel tag head tag_sha
channel=$(_lm_channel "config/config.json")
if [ "$channel" = "beta" ]; then
return 1
fi
if ! git fetch --quiet --tags --force origin >/dev/null 2>&1; then
print_warning "Could not fetch release tags; keeping the current version"
return 0
fi
tag=$(_lm_newest_release_tag)
head=$(git rev-parse --verify --quiet HEAD 2>/dev/null || true)
if [ -n "$tag" ] && [ -n "$head" ] && git merge-base --is-ancestor "$head" "$tag" 2>/dev/null; then
tag_sha=$(git rev-parse --verify --quiet "${tag}^{commit}" 2>/dev/null || true)
if [ "$head" = "$tag_sha" ]; then
print_success "Already on the newest release, $tag"
elif git -c advice.detachedHead=false checkout --quiet --detach "${tag}^{commit}"; then
print_success "Updated to release $tag (stable channel)"
else
print_warning "Could not move to release $tag (local changes?); keeping the current version"
fi
return 0
fi
if git symbolic-ref --quiet HEAD >/dev/null 2>&1; then
# Newer than the newest release (or no release yet): follow the branch
# until a release includes this version, as updates do.
return 1
fi
print_success "This checkout is newer than the newest release${tag:+ ($tag)}; leaving it as it is"
return 0
}
# --- end release checkout helpers --------------------------------------------
# Main installation function # Main installation function
main() { main() {
print_step "LED Matrix One-Shot Installation" print_step "LED Matrix One-Shot Installation"
@@ -292,7 +404,10 @@ main() {
# Try to safely update current branch first (fast-forward only to avoid unintended merges) # Try to safely update current branch first (fast-forward only to avoid unintended merges)
PULL_SUCCESS=false PULL_SUCCESS=false
if git pull --ff-only origin "$CURRENT_BRANCH" >/dev/null 2>&1; then # Stable: the newest release, if it contains this version.
if _lm_update_existing_checkout; then
PULL_SUCCESS=true
elif git pull --ff-only origin "$CURRENT_BRANCH" >/dev/null 2>&1; then
print_success "Repository updated successfully (branch: $CURRENT_BRANCH)" print_success "Repository updated successfully (branch: $CURRENT_BRANCH)"
PULL_SUCCESS=true PULL_SUCCESS=true
else else
@@ -323,10 +438,12 @@ main() {
rm -rf "$REPO_DIR" rm -rf "$REPO_DIR"
print_success "Cloning repository..." print_success "Cloning repository..."
retry git clone "$REPO_URL" "$REPO_DIR" retry git clone "$REPO_URL" "$REPO_DIR"
(cd "$REPO_DIR" && _lm_checkout_release_after_clone) || print_warning "Could not choose a release; installing the newest code from main"
fi fi
else else
print_success "Cloning repository to $REPO_DIR..." print_success "Cloning repository to $REPO_DIR..."
retry git clone "$REPO_URL" "$REPO_DIR" retry git clone "$REPO_URL" "$REPO_DIR"
(cd "$REPO_DIR" && _lm_checkout_release_after_clone) || print_warning "Could not choose a release; installing the newest code from main"
fi fi
# Verify repository is accessible # Verify repository is accessible
@@ -397,6 +514,7 @@ main() {
sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 \ sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 \
LEDMATRIX_APT_UPDATED="${LEDMATRIX_APT_UPDATED:-0}" \ LEDMATRIX_APT_UPDATED="${LEDMATRIX_APT_UPDATED:-0}" \
LEDMATRIX_AUTO_UPDATE="${LEDMATRIX_AUTO_UPDATE:-}" \ LEDMATRIX_AUTO_UPDATE="${LEDMATRIX_AUTO_UPDATE:-}" \
LEDMATRIX_CHANNEL="${LEDMATRIX_CHANNEL:-}" \
bash ./first_time_install.sh -y </dev/null bash ./first_time_install.sh -y </dev/null
fi fi
INSTALL_EXIT_CODE=$? INSTALL_EXIT_CODE=$?
+63 -19
View File
@@ -4,13 +4,14 @@ Alternative dependency installer that tries apt packages first,
then falls back to pip with --break-system-packages then falls back to pip with --break-system-packages
""" """
import re
import subprocess import subprocess
import sys import sys
import tempfile import tempfile
import warnings import warnings
from collections import deque from collections import deque
from pathlib import Path from pathlib import Path
from typing import List, Tuple from typing import Dict, List, Tuple
# How many trailing lines of a failed command's output to keep for the # How many trailing lines of a failed command's output to keep for the
# end-of-run failure summary. Keeps the root cause near the end of the log, # end-of-run failure summary. Keeps the root cause near the end of the log,
@@ -81,6 +82,8 @@ def install_via_pip(package_name: str) -> Tuple[bool, str]:
Returns (success, output). Returns (success, output).
""" """
# pip knows PIL as Pillow; the others are asked for by their own name.
package_name = _dist_name(package_name)
print(f"Installing {package_name} via pip...") print(f"Installing {package_name} via pip...")
success, output = _run([ success, output = _run([
sys.executable, '-m', 'pip', 'install', sys.executable, '-m', 'pip', 'install',
@@ -99,26 +102,66 @@ IMPORT_NAME_MAP = {
'freetype-py': 'freetype', 'freetype-py': 'freetype',
} }
# Minimum versions that must be met for an already-installed package to count # The packages above are keyed by what main() lists; these are the ones whose
# as satisfied. Debian Bookworm's python3-freetype is 2.3.0, below the # pip distribution name differs from that key.
# freetype-py>=2.5.1 pin in requirements.txt, so an import-only check would DIST_NAME_MAP = {
# wrongly skip the pip upgrade. 'PIL': 'Pillow',
MIN_VERSIONS = {
'freetype-py': (2, 5, 1),
} }
REQUIREMENTS_FILE = Path(__file__).resolve().parent.parent / 'web_interface' / 'requirements.txt'
def _version_tuple(text: str) -> tuple:
parts = []
for part in text.split('.'):
digits = ''.join(ch for ch in part if ch.isdigit())
if not digits:
break
parts.append(int(digits))
return tuple(parts)
def _requirement_floors(path: Path = REQUIREMENTS_FILE) -> Dict[str, tuple]:
"""``>=`` floors from a requirements file, keyed by lower-cased name.
The apt copies of these packages are older than the pins on both
supported releases -- Bookworm ships Flask and Werkzeug 2.2.2, Pillow 9.4,
requests 2.28, psutil 5.9, pytz 2022.7 and freetype-py 2.3; Trixie ships
Flask 3.1.1, Werkzeug 3.1.3, Pillow 11.1 and requests 2.32 --
so a package that merely imports is not enough. Read from the file rather
than copied here so the two cannot drift.
"""
floors: Dict[str, tuple] = {}
try:
lines = path.read_text(encoding='utf-8').splitlines()
except OSError:
return floors
for line in lines:
match = re.match(r'\s*([A-Za-z0-9][A-Za-z0-9._-]*)[^#]*?>=\s*([0-9][0-9.]*)', line)
if match:
floors[match.group(1).lower()] = _version_tuple(match.group(2))
return floors
def _dist_name(package_name: str) -> str:
return DIST_NAME_MAP.get(package_name, package_name)
def _minimum_version(package_name: str) -> tuple:
"""The required floor for ``package_name``, or () when there is none."""
return MIN_VERSIONS.get(_dist_name(package_name).lower(), ())
# Minimum versions that must be met for an already-installed package to count
# as satisfied.
MIN_VERSIONS = _requirement_floors()
def _installed_version_tuple(dist_name: str) -> tuple: def _installed_version_tuple(dist_name: str) -> tuple:
"""Return the installed distribution version as an int tuple, or () if unknown.""" """Return the installed distribution version as an int tuple, or () if unknown."""
try: try:
from importlib.metadata import version from importlib.metadata import version
parts = [] return _version_tuple(version(dist_name))
for part in version(dist_name).split('.'):
digits = ''.join(ch for ch in part if ch.isdigit())
if not digits:
break
parts.append(int(digits))
return tuple(parts)
except Exception: except Exception:
return () return ()
@@ -134,9 +177,9 @@ def check_package_installed(package_name: str) -> bool:
__import__(import_name) __import__(import_name)
except ImportError: except ImportError:
return False return False
minimum = MIN_VERSIONS.get(package_name) minimum = _minimum_version(package_name)
if minimum: if minimum:
installed = _installed_version_tuple(package_name) installed = _installed_version_tuple(_dist_name(package_name))
if not installed or installed < minimum: if not installed or installed < minimum:
print(f"{package_name} is installed but below the required " print(f"{package_name} is installed but below the required "
f"{'.'.join(map(str, minimum))}; will upgrade via pip") f"{'.'.join(map(str, minimum))}; will upgrade via pip")
@@ -188,10 +231,11 @@ def main():
continue continue
# Try apt first, then pip. An apt install only counts if it also # Try apt first, then pip. An apt install only counts if it also
# satisfies any minimum version (Debian's python3-freetype can be # satisfies the requirements floor (the apt copies of most of these
# older than the freetype-py pin), otherwise fall through to pip. # are older than the pins on both Bookworm and Trixie), otherwise
# fall through to pip.
ok, apt_output = install_via_apt(package) ok, apt_output = install_via_apt(package)
if ok and package in MIN_VERSIONS and not check_package_installed(package): if ok and _minimum_version(package) and not check_package_installed(package):
ok = False ok = False
apt_output = f"apt version of {package} is below the required minimum" apt_output = f"apt version of {package} is below the required minimum"
if not ok: if not ok:
+24 -4
View File
@@ -15,7 +15,8 @@ The updater leaves data/auto_update_pending.json:
{"status": "pending", "old_head": ..., "new_head": ..., {"status": "pending", "old_head": ..., "new_head": ...,
"old_ref": "main" | "" (detached) | absent (older updaters), "old_ref": "main" | "" (detached) | absent (older updaters),
"display_was_active": bool, "dependency_failures": [...], "created_at": ...} "display_was_active": bool, "dependency_failures": [...],
"units_refreshed": bool (absent from older updaters), "created_at": ...}
This moves its status to "verifying" and then to one of "success", This moves its status to "verifying" and then to one of "success",
"rolled_back" or "rollback_failed", with "reason" and "detail" saying why. "rolled_back" or "rollback_failed", with "reason" and "detail" saying why.
@@ -74,6 +75,10 @@ BASH_CANDIDATES = ('/usr/bin/bash', '/bin/bash')
#: ...and, like it, moves to the next one only when sudo refused the command #: ...and, like it, moves to the next one only when sudo refused the command
#: line (permission_utils.SUDO_REFUSAL_PHRASES), never after pip itself ran. #: line (permission_utils.SUDO_REFUSAL_PHRASES), never after pip itself ran.
SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no tty present') SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no tty present')
#: The root-owned helper that installed the update's systemd units
#: (web_interface/unit_refresh.py); ``--restore`` puts the previous ones back.
REFRESH_UNITS_PATH = '/usr/local/sbin/ledmatrix-refresh-units'
UNIT_RESTORE_TIMEOUT_SECONDS = 90
#: The longest one health check can take: restart and wait, roll back #: The longest one health check can take: restart and wait, roll back
#: (diff, reset, reinstalls), restart and wait again. A wait's last poll can #: (diff, reset, reinstalls), restart and wait again. A wait's last poll can
@@ -81,7 +86,8 @@ SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no t
_WAIT_WORST_SECONDS = (HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS + WEB_CHECK_TIMEOUT_SECONDS _WAIT_WORST_SECONDS = (HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS + WEB_CHECK_TIMEOUT_SECONDS
+ 2 * SYSTEMCTL_QUERY_TIMEOUT_SECONDS + POLL_SECONDS) + 2 * SYSTEMCTL_QUERY_TIMEOUT_SECONDS + POLL_SECONDS)
WORST_CASE_SECONDS = (2 * (2 * RESTART_TIMEOUT_SECONDS + _WAIT_WORST_SECONDS) WORST_CASE_SECONDS = (2 * (2 * RESTART_TIMEOUT_SECONDS + _WAIT_WORST_SECONDS)
+ GIT_TIMEOUT_SECONDS + GIT_RESET_TIMEOUT_SECONDS + PIP_BUDGET_SECONDS) + GIT_TIMEOUT_SECONDS + GIT_RESET_TIMEOUT_SECONDS + UNIT_RESTORE_TIMEOUT_SECONDS
+ PIP_BUDGET_SECONDS)
#: What a command that could not run at all reports: its callers only read #: What a command that could not run at all reports: its callers only read
#: these three fields, the same ones a completed subprocess has. #: these three fields, the same ones a completed subprocess has.
@@ -300,12 +306,26 @@ class Verifier:
if result.returncode != 0: if result.returncode != 0:
return False, (f'"git reset --hard {old}" failed: ' return False, (f'"git reset --hard {old}" failed: '
f'{(result.stderr or result.stdout or "").strip()}') f'{(result.stderr or result.stdout or "").strip()}')
notes = []
# The update also installed its own systemd units: put the previous
# ones back before anything restarts onto the rolled-back code.
if pending.get('units_refreshed') and not self.restore_units():
notes.append('restoring the previous service settings failed; run '
'"sudo ./scripts/install/install_service.sh" in the LEDMatrix folder')
deadline = self.clock() + PIP_BUDGET_SECONDS deadline = self.clock() + PIP_BUDGET_SECONDS
failed = [rel for rel in requirements if not self.install_requirements(rel, deadline)] failed = [rel for rel in requirements if not self.install_requirements(rel, deadline)]
if failed: if failed:
return True, ('reinstalling the previous dependencies from ' + ', '.join(failed) notes.append('reinstalling the previous dependencies from ' + ', '.join(failed)
+ ' failed; run Install Base Requirements from the Tools tab') + ' failed; run Install Base Requirements from the Tools tab')
return True, '' return True, '; '.join(notes)
def restore_units(self):
"""Reinstall the systemd units the update replaced. True on success."""
result = self._run(['sudo', '-n', REFRESH_UNITS_PATH, '--restore'],
timeout=UNIT_RESTORE_TIMEOUT_SECONDS)
if result.returncode != 0:
self.log(f'restoring the previous systemd units failed: {(result.stderr or "").strip()}')
return result.returncode == 0
# -- the check itself ------------------------------------------------- # -- the check itself -------------------------------------------------
+34 -12
View File
@@ -28,6 +28,7 @@ freetype.Face, so it drops straight into DisplayManager.draw_text().
""" """
import logging import logging
import weakref
from collections import OrderedDict from collections import OrderedDict
from dataclasses import dataclass from dataclasses import dataclass
from typing import Any, Dict, List, Optional, Sequence, Tuple, Union from typing import Any, Dict, List, Optional, Sequence, Tuple, Union
@@ -332,9 +333,10 @@ class LayoutContext:
# a plugin fitting changing text (a live game clock, a ticker) on a # a plugin fitting changing text (a live game clock, a ticker) on a
# 24/7 service would otherwise grow this without bound. # 24/7 service would otherwise grow this without bound.
self._fit_cache: "OrderedDict[Any, FitResult]" = OrderedDict() self._fit_cache: "OrderedDict[Any, FitResult]" = OrderedDict()
# LRU-bounded (images are big). Entries hold a strong reference to # LRU-bounded (images are big). An id()-keyed entry watches its
# the source image when keyed by id() so the id can't be recycled # source image through a weak reference and is dropped when the
# out from under the cache. # source is freed (see fit_image), so the id can't be recycled out
# from under the cache and the cache never keeps the source alive.
self._image_cache: "OrderedDict[Any, Tuple[Any, Any]]" = OrderedDict() self._image_cache: "OrderedDict[Any, Tuple[Any, Any]]" = OrderedDict()
_IMAGE_CACHE_MAX = 64 _IMAGE_CACHE_MAX = 64
@@ -536,8 +538,15 @@ class LayoutContext:
cached per (image, box size, options) for this panel size. cached per (image, box size, options) for this panel size.
Prefer a stable ``cache_key`` (e.g. "logo:KC") for images that get Prefer a stable ``cache_key`` (e.g. "logo:KC") for images that get
reloaded — the default id()-based key is safe (the entry pins the reloaded — the default id()-based key misses across reloads of the
source image) but misses across reloads of the same content. same content.
An id()-keyed entry lives only as long as its source image: it holds
a weak reference and is dropped when the source is freed. It used to
pin the source instead, so a plugin passing a freshly loaded image
each frame (``draw_image(Image.open(path), box)``, the documented
one-liner) never hit and kept the last 64 sources alive — ~64MB for
500x500 RGBA team logos, the median size under assets/sports.
""" """
from src.adaptive_images import fit_image as _fit_image from src.adaptive_images import fit_image as _fit_image
@@ -547,18 +556,31 @@ class LayoutContext:
key = ("image", identity, img.size, box_w, box_h, mode, key = ("image", identity, img.size, box_w, box_h, mode,
crop_to_ink, anchor, resample_name, upscale) crop_to_ink, anchor, resample_name, upscale)
cached = self._image_cache.get(key) cache = self._image_cache
if cached is not None: cached = cache.get(key)
self._image_cache.move_to_end(key) # An id()-keyed hit must still be this very image; the callback below
# normally removes a dead source's entry before its id can recur.
if cached is not None and (cache_key is not None or cached[1]() is img):
cache.move_to_end(key)
return cached[0] return cached[0]
result = _fit_image(img, (box_w, box_h), mode=mode, result = _fit_image(img, (box_w, box_h), mode=mode,
crop_to_ink=crop_to_ink, anchor=anchor, crop_to_ink=crop_to_ink, anchor=anchor,
resample=resample, upscale=upscale) resample=resample, upscale=upscale)
# Pin the source only for id()-keyed entries (see docstring). source = None
self._image_cache[key] = (result, img if cache_key is None else None) if cache_key is None:
while len(self._image_cache) > self._IMAGE_CACHE_MAX: def _forget(ref: Any, key: Any = key) -> None:
self._image_cache.popitem(last=False) entry = cache.get(key)
if entry is not None and entry[1] is ref:
cache.pop(key, None)
try:
source = weakref.ref(img, _forget)
except TypeError:
# Not weak-referenceable: pin it, as before.
source = lambda img=img: img # noqa: E731
cache[key] = (result, source)
while len(cache) > self._IMAGE_CACHE_MAX:
cache.popitem(last=False)
return result return result
# ---- text utilities ------------------------------------------------ # ---- text utilities ------------------------------------------------
+22 -13
View File
@@ -20,6 +20,7 @@ from typing import Dict, Any, Optional, List, cast
from src.common.api_helper import DEFAULT_HTTP_HEADERS from src.common.api_helper import DEFAULT_HTTP_HEADERS
from src.common.fetch_service import fetch_get, share_connection_pool from src.common.fetch_service import fetch_get, share_connection_pool
from src.common.json_body import response_json
@@ -146,7 +147,7 @@ class BaseOddsManager:
if _is_no_odds_marker(cached_data): if _is_no_odds_marker(cached_data):
self.logger.debug("Cached no-odds marker for %s", cache_key) self.logger.debug("Cached no-odds marker for %s", cache_key)
return None return None
self.logger.debug(f"Using cached odds from ESPN for {cache_key}") self.logger.debug("Using cached odds from ESPN for %s", cache_key)
return cached_data return cached_data
if time.monotonic() < self._skip_network_until: if time.monotonic() < self._skip_network_until:
@@ -159,7 +160,7 @@ class BaseOddsManager:
self._skip_network_until - time.monotonic()) self._skip_network_until - time.monotonic())
return None return None
self.logger.debug(f"Cache miss - fetching fresh odds from ESPN for {cache_key}") self.logger.debug("Cache miss - fetching fresh odds from ESPN for %s", cache_key)
try: try:
# Map league names to ESPN API format # Map league names to ESPN API format
@@ -173,26 +174,30 @@ class BaseOddsManager:
espn_league = league_mapping.get(league, league) espn_league = league_mapping.get(league, league)
url = f"{self.base_url}/{sport}/leagues/{espn_league}/events/{event_id}/competitions/{event_id}/odds" url = f"{self.base_url}/{sport}/leagues/{espn_league}/events/{event_id}/competitions/{event_id}/odds"
self.logger.debug(f"Requesting odds from URL: {url}") self.logger.debug("Requesting odds from URL: %s", url)
# The response cache may answer only inside this caller's own # The response cache may answer only inside this caller's own
# interval, the age at which its cached odds expire anyway. # interval, the age at which its cached odds expire anyway.
response = fetch_get(self.session, url, timeout=self.request_timeout, response = fetch_get(self.session, url, timeout=self.request_timeout,
cache_max_age=interval) cache_max_age=interval)
response.raise_for_status() response.raise_for_status()
raw_data = response.json() raw_data = response_json(response)
self._skip_network_until = 0.0 # reachable again self._skip_network_until = 0.0 # reachable again
self.logger.debug(f"Received raw odds data from ESPN: {json.dumps(raw_data, indent=2)}") # Guarded, not just %-style: the json.dumps argument would still be
# built for every response with DEBUG off.
if self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug("Received raw odds data from ESPN: %s",
json.dumps(raw_data, indent=2))
odds_data = self._extract_espn_data(raw_data) odds_data = self._extract_espn_data(raw_data)
if odds_data: if odds_data:
self.logger.debug(f"Successfully extracted odds data: {odds_data}") self.logger.debug("Successfully extracted odds data: %s", odds_data)
self.cache_manager.set(cache_key, odds_data, ttl=interval) self.cache_manager.set(cache_key, odds_data, ttl=interval)
self.logger.debug(f"Saved odds data to cache for {cache_key} with TTL {interval}s") self.logger.debug("Saved odds data to cache for %s with TTL %ss", cache_key, interval)
else: else:
self.logger.debug(f"No odds data available for {cache_key}") self.logger.debug("No odds data available for %s", cache_key)
# Cache the absence too, so the game is not re-requested # Cache the absence too, so the game is not re-requested
# on every update until the interval passes. # on every update until the interval passes.
self.cache_manager.set(cache_key, {"no_odds": True}, ttl=interval) self.cache_manager.set(cache_key, {"no_odds": True}, ttl=interval)
@@ -226,12 +231,12 @@ class BaseOddsManager:
Returns: Returns:
Formatted odds data dictionary or None Formatted odds data dictionary or None
""" """
self.logger.debug(f"Extracting ESPN odds data. Data keys: {list(data.keys())}") self.logger.debug("Extracting ESPN odds data. Data keys: %s", list(data.keys()))
if "items" in data and data["items"]: if "items" in data and data["items"]:
self.logger.debug(f"Found {len(data['items'])} items in odds data") self.logger.debug("Found %d items in odds data", len(data['items']))
item = data["items"][0] item = data["items"][0]
self.logger.debug(f"First item keys: {list(item.keys())}") self.logger.debug("First item keys: %s", list(item.keys()))
# The ESPN API returns odds data directly in the item, not in a # The ESPN API returns odds data directly in the item, not in a
# providers array. ESPN sends explicit JSON nulls for absent # providers array. ESPN sends explicit JSON nulls for absent
@@ -254,13 +259,17 @@ class BaseOddsManager:
.get("pointSpread") or {}).get("value") .get("pointSpread") or {}).get("value")
} }
} }
self.logger.debug(f"Returning extracted odds data: {json.dumps(extracted_data, indent=2)}") if self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug("Returning extracted odds data: %s",
json.dumps(extracted_data, indent=2))
return extracted_data return extracted_data
# Check if this is a valid empty response or an unexpected structure # Check if this is a valid empty response or an unexpected structure
if "count" in data and data["count"] == 0 and "items" in data and data["items"] == []: if "count" in data and data["count"] == 0 and "items" in data and data["items"] == []:
# This is a valid empty response - no odds available for this game # This is a valid empty response - no odds available for this game
self.logger.debug(f"No odds available for this game. Response: {json.dumps(data, indent=2)}") if self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug("No odds available for this game. Response: %s",
json.dumps(data, indent=2))
return None return None
else: else:
# This is an unexpected response structure # This is an unexpected response structure
+214 -35
View File
@@ -4,6 +4,7 @@ Disk Cache
Handles persistent disk-based caching with atomic writes and error recovery. Handles persistent disk-based caching with atomic writes and error recovery.
""" """
import hashlib
import json import json
import math import math
import os import os
@@ -14,7 +15,7 @@ import tempfile
import logging import logging
import threading import threading
import zlib import zlib
from typing import Dict, Any, Optional, Protocol from typing import Dict, Any, Optional, Protocol, Tuple
from datetime import datetime from datetime import datetime
from src.common.path_safety import safe_path_component from src.common.path_safety import safe_path_component
@@ -31,6 +32,35 @@ except ImportError: # pragma: no cover - exercised on hosts without the wheel
# useful, and a half-written file was never useful. # useful, and a half-written file was never useful.
_ORPHAN_TEMP_MAX_AGE_SECONDS = 3600 _ORPHAN_TEMP_MAX_AGE_SECONDS = 3600
# Longest key, in UTF-8 bytes, used verbatim as a filename stem. ext4 caps a
# name at 255 bytes and set()'s temp file is ".<stem>.json.<8 random>", 15
# bytes longer than the stem, so anything near the cap could never be written:
# the calendar plugin's key joins every calendar id and passed 300 bytes on a
# real install, failing every write with ENAMETOOLONG. Longer keys keep this
# many bytes as a readable prefix and end in a hash of the whole key.
_MAX_KEY_FILENAME_BYTES = 200
_KEY_HASH_CHARS = 16
def _filename_stem(key: str) -> str:
"""The filename stem for a key that is already a safe path component.
Short keys are used as they are, so every file already on disk keeps its
name. A long one becomes its first bytes plus a hash of the full key: the
prefix keeps the stem recognisable (and keeps the data-type words that
cleanup's retention lookup reads from it), the hash keeps two keys that
share a long prefix apart. The result is itself short, so a stem read back
from a filename -- which is how the web UI names a key it deletes -- maps to
the same file.
"""
encoded = key.encode('utf-8')
if len(encoded) <= _MAX_KEY_FILENAME_BYTES:
return key
digest = hashlib.sha256(encoded).hexdigest()[:_KEY_HASH_CHARS]
keep = _MAX_KEY_FILENAME_BYTES - _KEY_HASH_CHARS - 1
prefix = encoded[:keep].decode('utf-8', errors='ignore')
return f"{prefix}-{digest}"
class CacheStrategyProtocol(Protocol): class CacheStrategyProtocol(Protocol):
@@ -111,18 +141,91 @@ _HEAD_RE = re.compile(
) )
def _stale_from_head(head: bytes, max_age: Optional[int], now: float) -> bool: # UNCHANGED RE-SAVES: THE FILE'S MTIME CARRIES THE NEWER TIMESTAMP
# ----------------------------------------------------------------
# Plugins re-save unchanged API data every update cycle, and every one of
# those saves was a full rewrite on the SD card. DiskCache.set skips the write
# when the payload matches the last one it wrote for the key -- but
# CacheManager.set stamps each record with time.time(), so for set() the
# payload never matched and the skip never fired.
#
# The digest now leaves out a header-first record's timestamp, so an unchanged
# set() is skipped. What the skip must not do is make the record look older
# than it is: the timestamp inside the file is from the last real write, and
# a reader in another process (the web interface, with memory_ttl=0) or after
# a restart would call fresh data stale. So the newer timestamp goes where it
# costs no data write -- the file's mtime -- and readers take a record's age
# from the newer of the two. The invariant that makes that safe:
#
# a file's mtime is the timestamp of the newest record saved for its key
#
# real write mtime is set to the record's own timestamp, so a record saved
# with an old timestamp (data as of some earlier time) cannot
# borrow freshness from the moment it hit the disk
# skip mtime is set to the skipped record's timestamp -- exactly what
# a rewrite would have stored, without the rewrite
#
# Readers of the on-disk timestamp, all of which go through _effective_timestamp:
# DiskCache.get (the header check and the full parse; it also returns the
# record with 'timestamp' set to the effective value, so CacheManager.get's
# max_age path, the memory tier hydrated from disk, and any plugin reading
# record['timestamp'] all see it). Readers that use mtime alone already see the
# newer value: the retention sweep below, CacheManager.list_cache_files (the
# web UI's cache list). Nothing else opens cache files: web_interface and
# scripts reach them only through CacheManager.
#
# Something other than this class can also move an mtime forward -- a copy
# without -p, an rsync without -t, a `touch`. (backup_manager.py does not
# back up or restore the cache directory, so the in-tree restore cannot.) That
# must not make old data fresh, so the lift is bounded: a reader never takes
# the mtime as more than _MAX_TIMESTAMP_LIFT past the embedded timestamp, and
# set() rewrites the file for real once a skip would need more than that, so
# an honest lift never reaches the bound. A file copied a day after it was
# written therefore reads at most an hour fresher than its contents say, and a
# 30-second live-score record from yesterday stays stale. CacheManager.set
# records written before this change have mtime == write time == embedded
# timestamp, give or take the write itself, and read exactly as before; a
# file an older version wrote or touched later than its embedded timestamp
# says reads at most the same hour fresher, once, until it is next saved.
#: Longest a skipped write may stand in for a real one, and so the furthest a
#: file's mtime is ever trusted past the record's own timestamp. Unchanged data
#: is rewritten at least this often, at most once an hour per key instead of
#: once per update cycle.
_MAX_TIMESTAMP_LIFT = 3600.0
def _record_timestamp(value: Any) -> Optional[float]:
"""A record's timestamp as a finite float, or None if it has no usable one."""
if isinstance(value, bool) or not isinstance(value, (int, float)):
return None
value = float(value)
return value if math.isfinite(value) else None
def _effective_timestamp(embedded: float, mtime: Optional[float]) -> float:
"""When a record was last saved: its timestamp, or the file's mtime if a
later unchanged save moved that forward -- never by more than
_MAX_TIMESTAMP_LIFT. See "UNCHANGED RE-SAVES" above."""
if mtime is None:
return embedded
return max(embedded, min(mtime, embedded + _MAX_TIMESTAMP_LIFT))
def _stale_from_head(head: bytes, max_age: Optional[int], now: float,
mtime: Optional[float] = None) -> bool:
"""True when a record's header alone shows it has expired. """True when a record's header alone shows it has expired.
Mirrors the expiry rule in DiskCache.get: a per-entry ttl wins over the Mirrors the expiry rule in DiskCache.get: a per-entry ttl wins over the
caller's max_age, and no limit at all means never stale. False whenever the caller's max_age, and no limit at all means never stale. False whenever the
header cannot be read, so the full parse decides as it always did. header cannot be read, so the full parse decides as it always did. ``mtime``
is the file's, which may carry a newer save than the header does.
""" """
match = _HEAD_RE.match(head) match = _HEAD_RE.match(head)
if not match: if not match:
return False return False
try: try:
timestamp = float(match.group(1)) timestamp = _effective_timestamp(float(match.group(1)), mtime)
limit = max_age limit = max_age
if match.group(2) is not None: if match.group(2) is not None:
ttl = float(match.group(2)) ttl = float(match.group(2))
@@ -179,7 +282,7 @@ else:
# -------------------------------------------- # --------------------------------------------
# The display service runs as root and the web interface as the installing # The display service runs as root and the web interface as the installing
# user, and the web interface reads records only the display writes # user, and the web interface reads records only the display writes
# (display_current_state, display_on_demand_state, plugin_metrics:*). Files are # (display_current_state, display_on_demand_state, plugin_metrics_snapshot). Files are
# written 0660, so the web interface can read one only through its group. # written 0660, so the web interface can read one only through its group.
# #
# The installers rely on the directory's setgid bit to set that group. That is # The installers rely on the directory's setgid bit to set that group. That is
@@ -248,11 +351,14 @@ class DiskCache:
self.cache_dir = cache_dir self.cache_dir = cache_dir
self.logger = logger or logging.getLogger(__name__) self.logger = logger or logging.getLogger(__name__)
self._lock = threading.Lock() self._lock = threading.Lock()
# key -> adler32 of the last payload successfully written to the # key -> what set() last put at the primary cache path: the adler32 of
# primary cache path; lets set() skip rewriting identical data # the payload (less a header-first timestamp), the timestamp the file
# (per-process only — worst case another process rewrites, never # holds (None for records without one), and the file's inode and size.
# a missed write). Guarded by _lock. # Lets set() skip rewriting identical data. Per-process only, and the
self._write_digests: Dict[str, int] = {} # inode/size check means another process's write is never mistaken
# for ours -- worst case a redundant write, never a missed one.
# Guarded by _lock.
self._write_digests: Dict[str, Tuple[int, Optional[float], int, int]] = {}
def get_cache_path(self, key: str) -> Optional[str]: def get_cache_path(self, key: str) -> Optional[str]:
""" """
@@ -267,6 +373,8 @@ class DiskCache:
derives them), so rejecting anything with a path component turns derives them), so rejecting anything with a path component turns
away only inputs that could never have been written here. away only inputs that could never have been written here.
A key too long to be a filename is shortened by _filename_stem.
Args: Args:
key: Cache key key: Cache key
@@ -280,7 +388,7 @@ class DiskCache:
if safe_key is None: if safe_key is None:
self.logger.warning("Rejected unsafe cache key %r", key) self.logger.warning("Rejected unsafe cache key %r", key)
return None return None
return os.path.join(self.cache_dir, f"{safe_key}.json") return os.path.join(self.cache_dir, f"{_filename_stem(safe_key)}.json")
def get(self, key: str, max_age: Optional[int] = 300) -> Optional[Dict[str, Any]]: def get(self, key: str, max_age: Optional[int] = 300) -> Optional[Dict[str, Any]]:
""" """
@@ -301,31 +409,40 @@ class DiskCache:
try: try:
with self._lock: with self._lock:
with open(cache_path, 'rb') as f: with open(cache_path, 'rb') as f:
# The open file's mtime, not the path's: the file a skipped
# write touched is the one being read.
mtime = os.fstat(f.fileno()).st_mtime
# Decide staleness from the header before paying for the # Decide staleness from the header before paying for the
# parse. A stale read is the common case for the biggest # parse. A stale read is the common case for the biggest
# records (a season schedule is re-fetched when its cache # records (a season schedule is re-fetched when its cache
# expires), and parsing 53MB to throw it away held the GIL # expires), and parsing 53MB to throw it away held the GIL
# for ~1.8s -- a visible freeze on the panel. # for ~1.8s -- a visible freeze on the panel.
if _stale_from_head(f.read(_HEAD_BYTES), max_age, time.time()): if _stale_from_head(f.read(_HEAD_BYTES), max_age, time.time(), mtime):
return None return None
f.seek(0) f.seek(0)
record = _loads(f.read()) record = _loads(f.read())
# Determine record timestamp (prefer embedded, else file mtime) # Determine record timestamp: the embedded one, moved forward by a
# later unchanged save if there was one (see "UNCHANGED RE-SAVES"),
# else the file mtime.
record_ts = None record_ts = None
if isinstance(record, dict): if isinstance(record, dict):
record_ts = record.get('timestamp') record_ts = record.get('timestamp')
if record_ts is None: if record_ts is None:
try: record_ts = mtime
record_ts = os.path.getmtime(cache_path) else:
except OSError: embedded_ts = _record_timestamp(record_ts)
record_ts = None if embedded_ts is None:
if record_ts is not None:
try: try:
record_ts = float(record_ts) record_ts = float(record_ts)
except (TypeError, ValueError): except (TypeError, ValueError):
record_ts = None record_ts = None
else:
record_ts = _effective_timestamp(embedded_ts, mtime)
if record_ts != embedded_ts:
# Hand the record back as a rewrite would have left it,
# so callers that age it themselves agree with us.
record['timestamp'] = record_ts
now = time.time() now = time.time()
@@ -403,7 +520,12 @@ class DiskCache:
self.logger.warning("Cache data for key '%s' not serializable: %s", key, e) self.logger.warning("Cache data for key '%s' not serializable: %s", key, e)
return return
digest = zlib.adler32(payload) timestamp = _record_timestamp(data.get('timestamp')) if isinstance(data, dict) else None
# A header-first record (CacheManager.set's layout) is compared without
# its timestamp, which differs on every save; see "UNCHANGED RE-SAVES".
# Any other layout is compared whole, as before.
head = _HEAD_RE.match(payload) if timestamp is not None else None
digest = zlib.adler32(memoryview(payload)[head.end(1):] if head else payload)
try: try:
# Atomic write to avoid partial/corrupt files # Atomic write to avoid partial/corrupt files
@@ -411,16 +533,10 @@ class DiskCache:
# Skip the disk entirely when this exact payload was already # Skip the disk entirely when this exact payload was already
# written for this key (plugins re-save unchanged API data # written for this key (plugins re-save unchanged API data
# every update cycle — each write is real SD-card wear). # every update cycle — each write is real SD-card wear).
# Refresh the file mtime so records that rely on it for TTL # A metadata touch is journal-cheap compared to rewriting
# (no embedded 'timestamp') don't expire early; a metadata # the data.
# touch is journal-cheap compared to rewriting the data. if self._skip_unchanged(key, cache_path, digest, timestamp):
if self._write_digests.get(key) == digest:
try:
os.utime(cache_path, None)
return return
except OSError:
# File vanished or perms changed — fall through and write
self._write_digests.pop(key, None)
tmp_dir = os.path.dirname(cache_path) tmp_dir = os.path.dirname(cache_path)
# Try to create temp file in cache directory first # Try to create temp file in cache directory first
@@ -458,7 +574,7 @@ class DiskCache:
# opened it in between was refused. # opened it in between was refused.
_share_open_file(tmp_file.fileno(), _shared_group(tmp_dir)) _share_open_file(tmp_file.fileno(), _shared_group(tmp_dir))
os.replace(tmp_path, cache_path) os.replace(tmp_path, cache_path)
self._write_digests[key] = digest self._remember_write(key, cache_path, digest, timestamp)
finally: finally:
if os.path.exists(tmp_path): if os.path.exists(tmp_path):
try: try:
@@ -471,13 +587,13 @@ class DiskCache:
with open(cache_path, 'wb') as cache_file: with open(cache_path, 'wb') as cache_file:
cache_file.write(payload) cache_file.write(payload)
_share_open_file(cache_file.fileno(), _shared_group(tmp_dir)) _share_open_file(cache_file.fileno(), _shared_group(tmp_dir))
self._write_digests[key] = digest self._remember_write(key, cache_path, digest, timestamp)
self.logger.debug("Wrote cache for %s directly (non-atomic)", key) self.logger.debug("Wrote cache for %s directly (non-atomic)", key)
except (IOError, OSError, PermissionError) as write_error: except (IOError, OSError, PermissionError) as write_error:
# If direct write also fails, try fallback location # If direct write also fails, try fallback location
self.logger.warning("Direct write failed for key '%s' to %s: %s", key, cache_path, write_error) self.logger.warning("Direct write failed for key '%s' to %s: %s", key, cache_path, write_error)
raise # Re-raise to trigger fallback logic raise # Re-raise to trigger fallback logic
except (IOError, OSError, PermissionError): except (IOError, OSError, PermissionError) as primary_error:
# Attempt one-time fallback write to user's home cache directory # Attempt one-time fallback write to user's home cache directory
try: try:
# Try user's home cache directory as fallback # Try user's home cache directory as fallback
@@ -503,11 +619,14 @@ class DiskCache:
self.logger.debug("Fallback cache write also failed for key '%s': %s", key, e2) self.logger.debug("Fallback cache write also failed for key '%s': %s", key, e2)
# If all write attempts failed, log warning but don't raise exception # If all write attempts failed, log warning but don't raise exception
# Cache is a performance optimization, not critical for operation # Cache is a performance optimization, not critical for operation.
# Name the real error: this used to say "permission denied"
# whatever happened, which sent a too-long filename off to
# be debugged as a directory-ownership problem.
self.logger.warning( self.logger.warning(
"Could not write cache for key '%s' to %s (permission denied). " "Could not write cache for key '%s' to %s (%s). "
"Cache will be unavailable for this key, but application will continue.", "Cache will be unavailable for this key, but application will continue.",
key, cache_path key, cache_path, primary_error.strerror or primary_error
) )
return # Exit gracefully without raising exception return # Exit gracefully without raising exception
@@ -520,6 +639,66 @@ class DiskCache:
) )
return # Exit gracefully without raising exception return # Exit gracefully without raising exception
def _skip_unchanged(self, key: str, cache_path: str, digest: int,
timestamp: Optional[float]) -> bool:
"""Stand in for a write of an unchanged record by touching the file.
True when the file already holds this record bar its timestamp and the
touch landed; False means write it. The touch sets mtime to the
record's timestamp -- what a rewrite would have stored -- or to now
for a record without one, whose age readers already take from mtime.
Caller holds _lock.
"""
last = self._write_digests.get(key)
if last is None or last[0] != digest:
return False
_, written_ts, ino, size = last
if (timestamp is None) != (written_ts is None):
return False
if timestamp is not None and written_ts is not None:
# Never backwards (a rewrite would make the record older), and
# never further than readers will trust the mtime: past that the
# record is rewritten, so its own timestamp catches up.
if not written_ts <= timestamp <= written_ts + _MAX_TIMESTAMP_LIFT:
return False
try:
st = os.stat(cache_path)
if (st.st_ino, st.st_size) != (ino, size):
# Replaced since our write (another process, a restore):
# its contents are not the ones the digest describes.
self._write_digests.pop(key, None)
return False
# Setting an explicit time needs the file's owner; a file someone
# else wrote fails here and is rewritten (as our own file) instead.
os.utime(cache_path, None if timestamp is None else (timestamp, timestamp))
return True
except OSError:
# File vanished or perms changed — fall through and write
self._write_digests.pop(key, None)
return False
def _remember_write(self, key: str, cache_path: str, digest: int,
timestamp: Optional[float]) -> None:
"""After a real write: pin mtime to the record's timestamp and note
what was written, so the next unchanged save can be skipped.
Pinning keeps a record saved with an older timestamp from looking as
fresh as the moment it was written (see "UNCHANGED RE-SAVES"); for
CacheManager.set's records the two differ only by the write itself.
A timestamp in the future is left alone, mtime already being older.
Never raises: the data is on disk, and anything failing here only
costs the next save its skip. Caller holds _lock.
"""
self._write_digests.pop(key, None)
try:
if timestamp is not None and timestamp <= time.time():
os.utime(cache_path, (timestamp, timestamp))
st = os.stat(cache_path)
except OSError as e:
self.logger.debug("Could not pin mtime of %s: %s", cache_path, e)
return
self._write_digests[key] = (digest, timestamp, st.st_ino, st.st_size)
def clear(self, key: Optional[str] = None) -> None: def clear(self, key: Optional[str] = None) -> None:
""" """
Clear cache entry or all entries. Clear cache entry or all entries.
+80 -9
View File
@@ -43,6 +43,35 @@ from src.logging_config import get_logger
# it from either path. # it from either path.
from src.cache.disk_cache import DateTimeEncoder # noqa: F401 - deliberate re-export from src.cache.disk_cache import DateTimeEncoder # noqa: F401 - deliberate re-export
# CacheManager.config_manager not built yet (None means "not available").
_UNSET: Any = object()
def _outlived(record: Any, max_age: Optional[float], now: float) -> bool:
"""Whether a record's own timestamp puts it past max_age.
The memory tier times an entry from when it was put there, and a record
loaded from disk is put there when it is read, not when it was written: a
record 290 s old, read after a restart, could be served for another
max_age from memory. This is the age check DiskCache.get makes, with the
same rule that a stored ttl wins over the caller's max_age. A record that
carries no timestamp is left to the memory tier's own clock.
"""
if not isinstance(record, dict):
return False
stored_ttl = record.get('ttl')
if isinstance(stored_ttl, (int, float)) and not isinstance(stored_ttl, bool) \
and stored_ttl >= 0:
max_age = stored_ttl
stamp = record.get('timestamp')
if max_age is None or stamp is None or isinstance(stamp, bool):
return False
try:
return now - float(stamp) > max_age
except (TypeError, ValueError):
return False
class CacheManager: class CacheManager:
"""Manages caching of API responses to reduce API calls.""" """Manages caching of API responses to reduce API calls."""
@@ -73,21 +102,19 @@ class CacheManager:
self.logger.error("Could not find or create a writable cache directory. Caching will be disabled.") self.logger.error("Could not find or create a writable cache directory. Caching will be disabled.")
self.cache_dir = None self.cache_dir = None
# Initialize config manager for sport-specific intervals # The config manager is built on first use of self.config_manager; see
try: # the property. Nothing in the cache reads it any more.
from src.config_manager import ConfigManager self._config_manager: Any = _UNSET
self.config_manager: Optional[Any] = ConfigManager() self._config_manager_lock = threading.Lock()
self.config_manager.load_config()
except ImportError:
self.config_manager: Optional[Any] = None
self.logger.warning("ConfigManager not available, using default cache intervals")
# Initialize cache components using composition # Initialize cache components using composition
self._memory_cache_component = MemoryCache( self._memory_cache_component = MemoryCache(
max_size=default_max_size(), cleanup_interval=300.0 max_size=default_max_size(), cleanup_interval=300.0
) )
self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger) self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger)
self._strategy_component = CacheStrategy(config_manager=self.config_manager, logger=self.logger) # No config manager: CacheStrategy keeps the parameter for callers but
# reads nothing from it, and passing ours would build it eagerly.
self._strategy_component = CacheStrategy(logger=self.logger)
self._metrics_component = CacheMetrics(logger=self.logger) self._metrics_component = CacheMetrics(logger=self.logger)
# Disk cleanup configuration # Disk cleanup configuration
@@ -115,6 +142,44 @@ class CacheManager:
if self.cache_dir: if self.cache_dir:
self.start_cleanup_thread() self.start_cleanup_thread()
@property
def config_manager(self) -> Optional[Any]:
"""A loaded ConfigManager, built the first time it is asked for.
Every CacheManager used to build one and load the whole config in
__init__, for a cache strategy that stopped reading it -- startup paid
a config load (and the web interface another) per manager for nothing.
It is still public: the sports plugins resolve the global timezone and
display settings through ``cache_manager.config_manager``, and they get
the same object they always did, on first access instead of at
construction. None when ConfigManager cannot be imported, as before.
Assigning replaces it, as assigning the attribute always did.
"""
# getattr: a manager made with __new__ (some tests) has no slot yet.
value = getattr(self, '_config_manager', _UNSET)
if value is not _UNSET:
return value
lock = getattr(self, '_config_manager_lock', None) or threading.Lock()
with lock:
value = getattr(self, '_config_manager', _UNSET)
if value is _UNSET:
try:
from src.config_manager import ConfigManager
except ImportError:
self.logger.warning("ConfigManager not available, using default cache intervals")
value = None
else:
value = ConfigManager()
# Raises as it did from __init__; nothing is kept, so the
# next access tries again.
value.load_config()
self._config_manager = value
return value
@config_manager.setter
def config_manager(self, value: Optional[Any]) -> None:
self._config_manager = value
def _get_writable_cache_dir(self) -> Optional[str]: def _get_writable_cache_dir(self) -> Optional[str]:
"""Tries to find or create a writable cache directory, preferring a system path when available.""" """Tries to find or create a writable cache directory, preferring a system path when available."""
# Attempt 1: System-wide persistent cache directory (preferred for services) # Attempt 1: System-wide persistent cache directory (preferred for services)
@@ -245,7 +310,11 @@ class CacheManager:
# 1) Memory cache # 1) Memory cache
cached = self._memory_cache_component.get(key, max_age=in_memory_ttl) cached = self._memory_cache_component.get(key, max_age=in_memory_ttl)
if cached is not None: if cached is not None:
if not _outlived(cached, max_age, time.time()):
return cached return cached
# Too old for this reader. Disk may hold a newer write (from the
# other process), and if it does not, the miss is the right answer.
self._memory_cache_component.clear(key)
# 2) Disk cache # 2) Disk cache
record = self._disk_cache_component.get(key, max_age=max_age) record = self._disk_cache_component.get(key, max_age=max_age)
@@ -279,7 +348,9 @@ class CacheManager:
# Check memory cache first (1 minute TTL) # Check memory cache first (1 minute TTL)
cached = self._memory_cache_component.get(key, max_age=60) cached = self._memory_cache_component.get(key, max_age=60)
if cached is not None: if cached is not None:
if not _outlived(cached, 3600, time.time()):
return cached return cached
self._memory_cache_component.clear(key)
# Check disk cache # Check disk cache
data = self._disk_cache_component.get(key, max_age=3600) # 1 hour for load_cache data = self._disk_cache_component.get(key, max_age=3600) # 1 hour for load_cache
+3 -1
View File
@@ -17,7 +17,9 @@ Rules for the package:
- `from src.common import ...` re-exports `APIHelper`, `ScrollHelper`, - `from src.common import ...` re-exports `APIHelper`, `ScrollHelper`,
`LogoHelper`, `TextHelper`, `scroll_config` (plus `ScrollSettings`, `LogoHelper`, `TextHelper`, `scroll_config` (plus `ScrollSettings`,
`configure_scroll`, `resolve_scroll_settings`, `refresh_hz_from_config`) and `configure_scroll`, `resolve_scroll_settings`, `refresh_hz_from_config`) and
the adaptive layout names below ([`__init__.py`](__init__.py)). the adaptive layout names below ([`__init__.py`](__init__.py)). Each is
imported on first use, so `import src.common` or a submodule import stays
cheap; add a new re-export to `_LAZY` there as well as `__all__`.
## Summary ## Summary
+82 -15
View File
@@ -6,25 +6,37 @@ This package provides reusable functionality for plugins and core modules:
- Logo helpers - Logo helpers
- Text/scroll helpers - Text/scroll helpers
- Adaptive layout and image helpers - Adaptive layout and image helpers
The names below are imported on first use (PEP 562), not when the package is
imported. ``from src.common import ScrollHelper`` and
``src.common.ScrollHelper`` work as before and return the same objects, but
``import src.common`` -- or importing any submodule, such as
``src.common.path_safety`` -- no longer loads numpy, requests and freetype
along with every helper. The web interface imports src.common only for a few
small modules and never needs those.
""" """
# Export commonly used utilities import importlib
from src.common.api_helper import APIHelper from typing import TYPE_CHECKING, Any, Dict, List, Optional, Tuple
from src.common.scroll_helper import ScrollHelper
from src.common import scroll_config if TYPE_CHECKING:
from src.common.scroll_config import ( # What mypy and editors see: the real names and their types.
from src.common.api_helper import APIHelper
from src.common.scroll_helper import ScrollHelper
from src.common import scroll_config
from src.common.scroll_config import (
ScrollSettings, ScrollSettings,
configure as configure_scroll, configure as configure_scroll,
resolve as resolve_scroll_settings, resolve as resolve_scroll_settings,
refresh_hz_from_config, refresh_hz_from_config,
) )
from src.common.logo_helper import LogoHelper from src.common.logo_helper import LogoHelper
from src.common.text_helper import TextHelper from src.common.text_helper import TextHelper
# Adaptive layout & images (canonical homes: src.adaptive_layout / # Adaptive layout & images (canonical homes: src.adaptive_layout /
# src.adaptive_images — re-exported here so plugin authors find them in the # src.adaptive_images — re-exported here so plugin authors find them in the
# blessed-helpers package). See docs/ADAPTIVE_LAYOUT.md. # blessed-helpers package). See docs/ADAPTIVE_LAYOUT.md.
from src.adaptive_layout import ( from src.adaptive_layout import (
Region, Region,
LayoutContext, LayoutContext,
FontStep, FontStep,
@@ -37,14 +49,47 @@ from src.adaptive_layout import (
scoreboard_regions, scoreboard_regions,
MediaRow, MediaRow,
media_row, media_row,
) )
from src.adaptive_images import ( from src.adaptive_images import (
ImageFitResult, ImageFitResult,
fit_image, fit_image,
draw_fitted_image, draw_fitted_image,
RESAMPLE_LANCZOS, RESAMPLE_LANCZOS,
RESAMPLE_NEAREST, RESAMPLE_NEAREST,
) )
#: Exported name -> (module it lives in, attribute name there). An attribute
#: of None means the name is the module itself. Keep in step with the
#: TYPE_CHECKING imports above and with __all__.
_LAZY: Dict[str, Tuple[str, Optional[str]]] = {
'APIHelper': ('src.common.api_helper', 'APIHelper'),
'ScrollHelper': ('src.common.scroll_helper', 'ScrollHelper'),
'scroll_config': ('src.common.scroll_config', None),
'ScrollSettings': ('src.common.scroll_config', 'ScrollSettings'),
'configure_scroll': ('src.common.scroll_config', 'configure'),
'resolve_scroll_settings': ('src.common.scroll_config', 'resolve'),
'refresh_hz_from_config': ('src.common.scroll_config', 'refresh_hz_from_config'),
'LogoHelper': ('src.common.logo_helper', 'LogoHelper'),
'TextHelper': ('src.common.text_helper', 'TextHelper'),
# adaptive layout & images
'Region': ('src.adaptive_layout', 'Region'),
'LayoutContext': ('src.adaptive_layout', 'LayoutContext'),
'FontStep': ('src.adaptive_layout', 'FontStep'),
'FontLadder': ('src.adaptive_layout', 'FontLadder'),
'LADDER_GRID': ('src.adaptive_layout', 'LADDER_GRID'),
'LADDER_ARCADE': ('src.adaptive_layout', 'LADDER_ARCADE'),
'FitResult': ('src.adaptive_layout', 'FitResult'),
'draw_fitted_text': ('src.adaptive_layout', 'draw_fitted_text'),
'ScoreboardRegions': ('src.adaptive_layout', 'ScoreboardRegions'),
'scoreboard_regions': ('src.adaptive_layout', 'scoreboard_regions'),
'MediaRow': ('src.adaptive_layout', 'MediaRow'),
'media_row': ('src.adaptive_layout', 'media_row'),
'ImageFitResult': ('src.adaptive_images', 'ImageFitResult'),
'fit_image': ('src.adaptive_images', 'fit_image'),
'draw_fitted_image': ('src.adaptive_images', 'draw_fitted_image'),
'RESAMPLE_LANCZOS': ('src.adaptive_images', 'RESAMPLE_LANCZOS'),
'RESAMPLE_NEAREST': ('src.adaptive_images', 'RESAMPLE_NEAREST'),
}
__all__ = [ __all__ = [
'APIHelper', 'APIHelper',
@@ -75,3 +120,25 @@ __all__ = [
'RESAMPLE_LANCZOS', 'RESAMPLE_LANCZOS',
'RESAMPLE_NEAREST', 'RESAMPLE_NEAREST',
] ]
def __getattr__(name: str) -> Any:
"""Import an exported name on first access (PEP 562).
Only called for names not already in the module namespace, so after the
first access the cached value below is returned directly. Unknown names
raise AttributeError, which ``from src.common import <submodule>`` relies
on to fall through to importing the submodule.
"""
try:
module_name, attr = _LAZY[name]
except KeyError:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None
module = importlib.import_module(module_name) # nosemgrep: python.lang.security.audit.non-literal-import.non-literal-import -- module_name comes from the fixed _LAZY table
value = module if attr is None else getattr(module, attr)
globals()[name] = value
return value
def __dir__() -> List[str]:
return sorted(set(globals()) | set(__all__))
+3 -2
View File
@@ -17,6 +17,7 @@ from src.common.espn_dates import (
store_espn_scoreboard_cache, store_espn_scoreboard_cache,
) )
from src.common.fetch_service import fetch_get, fetch_post, share_connection_pool from src.common.fetch_service import fetch_get, fetch_post, share_connection_pool
from src.common.json_body import response_json
from typing import TYPE_CHECKING, Any, Dict, Mapping, Optional, cast from typing import TYPE_CHECKING, Any, Dict, Mapping, Optional, cast
import requests import requests
@@ -157,7 +158,7 @@ class APIHelper:
response.raise_for_status() response.raise_for_status()
# Parse JSON response # Parse JSON response
data: Dict[Any, Any] = response.json() data: Dict[Any, Any] = response_json(response)
# Cache response if cache key provided # Cache response if cache key provided
if cache_key and self.cache_manager: if cache_key and self.cache_manager:
@@ -304,7 +305,7 @@ class APIHelper:
) )
response.raise_for_status() response.raise_for_status()
return cast(Optional[Dict[Any, Any]], response.json()) return cast(Optional[Dict[Any, Any]], response_json(response))
except requests.exceptions.RequestException as e: except requests.exceptions.RequestException as e:
self.logger.error(f"POST request failed for {url}: {e}") self.logger.error(f"POST request failed for {url}: {e}")
+145 -2
View File
@@ -46,19 +46,39 @@ entry older than the reader's own ``max_age``, whoever wrote it and whatever
ttl they stored with it. Old keys are passed as ``legacy_keys`` and read ttl they stored with it. Old keys are passed as ``legacy_keys`` and read
after the canonical one, so an upgrade does not refetch everything at once; after the canonical one, so an upgrade does not refetch everything at once;
they can go one release after the one that added this. they can go one release after the one that added this.
Chunks whose days are long over are kept in memory between fetches. The
scoreboards re-fetch their whole Recent/Upcoming window (14 days back, 7
ahead) every hour, and since ranges went away that is 22 day requests per
league. Measured on hdpi on 2026-10-02 (NFL, college football, MLB, college
baseball, NHL): the hourly window refresh was ~270 of 321 ESPN requests and
~21 of 24.6MB in the hour, and the 12 days that ended three or more days ago
were 68% of those bytes (6.9 of 10.2MB per copy of the five windows). A
settled chunk is answered from memory for ``SETTLED_CHUNK_TTL_SECONDS``,
stored as zlib-compressed JSON (~13x smaller than the body, and far smaller
than the parsed objects), so the hourly refresh only goes to ESPN for the
days that can still change.
""" """
import contextvars import contextvars
import json
import logging import logging
import math import math
import re import re
import threading import threading
import time import time
import zlib
from collections import OrderedDict
from concurrent.futures import ThreadPoolExecutor from concurrent.futures import ThreadPoolExecutor
from datetime import date, datetime, timedelta from datetime import date, datetime, timedelta, timezone
from functools import partial from functools import partial
from typing import Any, Callable, Dict, Iterable, List, Optional, Tuple, cast from typing import Any, Callable, Dict, Iterable, List, Optional, Tuple, cast
try:
import orjson
except ImportError: # optional; the stdlib parser gives the same objects
orjson = None
try: try:
from src.common.json_body import response_json from src.common.json_body import response_json
except ImportError: except ImportError:
@@ -103,11 +123,37 @@ ESPN_CHUNK_WORKERS = 6
_range_lock = threading.Lock() _range_lock = threading.Lock()
_ranges_rejected_until = 0.0 _ranges_rejected_until = 0.0
# A chunk is "settled" once its last day is this many UTC days back. ESPN
# files games under the US Eastern date, and a late West-coast game ends after
# midnight UTC; three days leaves a full day of margin past both, so nothing
# still being played, finalised or rescheduled is ever served from memory.
SETTLED_AFTER_DAYS = 3
# How long a settled chunk is trusted. A day's finals do not change, but a
# rare correction (or an empty answer during an ESPN outage) should not live
# forever: once a day is plenty, and still skips 23 of every 24 hourly asks.
SETTLED_CHUNK_TTL_SECONDS = 24 * 60 * 60
# Bounds on the settled-chunk memory. A settled day measured 90KB (NHL) to
# 990KB (a college-football Saturday) of JSON and 9-74KB compressed; the five
# windows on hdpi need 60 entries and ~0.55MB. The caps only matter for a
# board fetching whole past seasons.
SETTLED_CACHE_MAX_ENTRIES = 512
SETTLED_CACHE_MAX_BYTES = 8 * 1024 * 1024
_settled_lock = threading.Lock()
# key -> (stored_at monotonic, compressed JSON)
_settled_chunks: "OrderedDict[Any, Tuple[float, bytes]]" = OrderedDict()
_settled_bytes = 0
__all__ = [ __all__ = [
"ESPN_MAX_LIMIT", "ESPN_MAX_LIMIT",
"ESPN_CHUNK_WORKERS", "ESPN_CHUNK_WORKERS",
"RANGE_RETRY_SECONDS", "RANGE_RETRY_SECONDS",
"SETTLED_AFTER_DAYS",
"SETTLED_CHUNK_TTL_SECONDS",
"clamp_espn_limit", "clamp_espn_limit",
"clear_settled_chunk_cache",
"parse_espn_date_range", "parse_espn_date_range",
"espn_date_chunks", "espn_date_chunks",
"merge_scoreboard_payloads", "merge_scoreboard_payloads",
@@ -199,6 +245,88 @@ def _days_of_month(chunk: str) -> List[str]:
return days return days
def _utc_today() -> date:
return datetime.now(timezone.utc).date()
def _chunk_last_day(chunk: str) -> Optional[date]:
try:
if len(chunk) == 8:
return date(int(chunk[:4]), int(chunk[4:6]), int(chunk[6:]))
if len(chunk) == 6:
first = date(int(chunk[:4]), int(chunk[4:6]), 1)
return _first_of_next_month(first) - timedelta(days=1)
except ValueError:
pass
return None
def _settled_key(url: str, params: Dict[str, Any], chunk: str) -> Optional[Any]:
"""Memory key for a chunk that can no longer change, else None."""
last_day = _chunk_last_day(chunk)
if last_day is None:
return None
if last_day > _utc_today() - timedelta(days=SETTLED_AFTER_DAYS):
return None
# dates is the chunk itself and limit is always ESPN_MAX_LIMIT here;
# anything else (groups=80 for FBS, a team filter) changes the answer.
rest = tuple(sorted(
(str(k), str(v)) for k, v in params.items() if k not in ("dates", "limit")
))
return (url, rest, chunk)
def _settled_get(key: Any) -> Optional[Dict[str, Any]]:
with _settled_lock:
entry = _settled_chunks.get(key)
if entry is None:
return None
if time.monotonic() - entry[0] > SETTLED_CHUNK_TTL_SECONDS:
_settled_drop(key)
return None
_settled_chunks.move_to_end(key)
blob = entry[1]
# Decompress and parse outside the lock: every hit gets its own objects,
# so a caller mutating its payload cannot reach another caller's.
body = zlib.decompress(blob)
return cast(Dict[str, Any], orjson.loads(body) if orjson else json.loads(body))
def _settled_drop(key: Any) -> None:
"""Remove one entry. Caller holds _settled_lock."""
global _settled_bytes
entry = _settled_chunks.pop(key, None)
if entry is not None:
_settled_bytes -= len(entry[1])
def _settled_put(key: Any, response: Any, payload: Dict[str, Any]) -> None:
global _settled_bytes
body = getattr(response, "content", None)
if not isinstance(body, (bytes, bytearray)):
body = json.dumps(payload).encode("utf-8")
blob = zlib.compress(bytes(body), 6)
if len(blob) > SETTLED_CACHE_MAX_BYTES:
return
with _settled_lock:
_settled_drop(key)
_settled_chunks[key] = (time.monotonic(), blob)
_settled_bytes += len(blob)
while _settled_chunks and (
len(_settled_chunks) > SETTLED_CACHE_MAX_ENTRIES
or _settled_bytes > SETTLED_CACHE_MAX_BYTES
):
_settled_drop(next(iter(_settled_chunks)))
def clear_settled_chunk_cache() -> None:
"""Forget every remembered settled chunk (tests, or a manual refresh)."""
global _settled_bytes
with _settled_lock:
_settled_chunks.clear()
_settled_bytes = 0
def espn_date_chunks(start: date, end: date) -> List[str]: def espn_date_chunks(start: date, end: date) -> List[str]:
"""Cover ``[start, end]`` inclusive with ``dates=`` values ESPN accepts. """Cover ``[start, end]`` inclusive with ``dates=`` values ESPN accepts.
@@ -255,8 +383,16 @@ def _fetch_one_chunk(
One bad chunk must not sink the rest of the season, so every error is One bad chunk must not sink the rest of the season, so every error is
logged and swallowed here rather than raised to the gather below. logged and swallowed here rather than raised to the gather below.
A chunk whose days are settled (see ``SETTLED_AFTER_DAYS``) is answered
from memory when it was fetched in the last day.
""" """
try: try:
settled = _settled_key(url, params, chunk)
if settled is not None:
cached = _settled_get(settled)
if cached is not None:
return cached
response = fetch_get( response = fetch_get(
session, session,
url, url,
@@ -266,7 +402,14 @@ def _fetch_one_chunk(
**_memo_kwargs(cache_max_age), **_memo_kwargs(cache_max_age),
) )
response.raise_for_status() response.raise_for_status()
return cast(Optional[Dict[str, Any]], response_json(response)) payload = response_json(response)
if settled is not None and isinstance(payload, dict):
events = payload.get("events")
# A capped month is truncated and gets re-asked day by day;
# remembering it would only cost memory.
if isinstance(events, list) and len(events) < ESPN_MAX_LIMIT:
_settled_put(settled, response, payload)
return cast(Optional[Dict[str, Any]], payload)
except Exception as exc: # noqa: BLE001 - see docstring except Exception as exc: # noqa: BLE001 - see docstring
if logger: if logger:
logger.warning("ESPN chunk %s failed, skipping it: %s", chunk, exc) logger.warning("ESPN chunk %s failed, skipping it: %s", chunk, exc)
+38 -3
View File
@@ -121,6 +121,7 @@ three times per threshold, so keep it to diagnostic runs, not soaks.
from __future__ import annotations from __future__ import annotations
import atexit
import copy import copy
import gc import gc
import json import json
@@ -213,10 +214,19 @@ class GcMonitor:
lock: the render thread and the stats writer only read them. lock: the render thread and the stats writer only read them.
Install it once per process with :func:`install_gc_monitor`. Install it once per process with :func:`install_gc_monitor`.
Collections still run while the interpreter shuts down, after module
globals such as ``time`` may already be torn down to ``None``. The clock
and ``sys.is_finalizing`` are bound here so the callback never looks a
global up, it does nothing once finalization has begun, and
:func:`install_gc_monitor` unregisters it at exit anyway.
""" """
def __init__(self, threshold: float = GC_PAUSE_SECONDS): def __init__(self, threshold: float = GC_PAUSE_SECONDS,
clock: Callable[[], float] = time.perf_counter):
self.threshold = threshold self.threshold = threshold
self._clock = clock
self._is_finalizing = sys.is_finalizing
self._started: Optional[float] = None self._started: Optional[float] = None
#: Per generation (0, 1, 2), since the monitor was installed. #: Per generation (0, 1, 2), since the monitor was installed.
self.collections = [0, 0, 0] self.collections = [0, 0, 0]
@@ -232,7 +242,9 @@ class GcMonitor:
self.last_long: Optional[Tuple[float, float]] = None self.last_long: Optional[Tuple[float, float]] = None
def __call__(self, phase: str, info: Dict[str, Any]) -> None: def __call__(self, phase: str, info: Dict[str, Any]) -> None:
now = time.perf_counter() if self._is_finalizing():
return
now = self._clock()
if phase == "start": if phase == "start":
self._started = now self._started = now
return return
@@ -267,15 +279,38 @@ _gc_monitor_lock = threading.Lock()
def install_gc_monitor() -> GcMonitor: def install_gc_monitor() -> GcMonitor:
"""The process's GcMonitor, installed in ``gc.callbacks`` on first call.""" """The process's GcMonitor, installed in ``gc.callbacks`` on first call.
It is unregistered at exit (:func:`uninstall_gc_monitor`), before the
interpreter tears module globals down.
"""
global _gc_monitor global _gc_monitor
with _gc_monitor_lock: with _gc_monitor_lock:
if _gc_monitor is None: if _gc_monitor is None:
_gc_monitor = GcMonitor() _gc_monitor = GcMonitor()
gc.callbacks.append(_gc_monitor) gc.callbacks.append(_gc_monitor)
atexit.register(uninstall_gc_monitor)
return _gc_monitor return _gc_monitor
def uninstall_gc_monitor() -> None:
"""Take the process's GcMonitor out of ``gc.callbacks``; safe to repeat.
A recorder that still holds the monitor keeps its counters; they just
stop moving. The next :func:`install_gc_monitor` installs a fresh one.
"""
global _gc_monitor
with _gc_monitor_lock:
monitor, _gc_monitor = _gc_monitor, None
if monitor is None:
return
atexit.unregister(uninstall_gc_monitor)
try:
gc.callbacks.remove(monitor)
except ValueError:
pass
#: One presented frame's interval: (interval, blit, wait, hold, ops), where #: One presented frame's interval: (interval, blit, wait, hold, ops), where
#: ops is the work noted before it (kind -> bytes) or None. #: ops is the work noted before it (kind -> bytes) or None.
_Frame = Tuple[float, float, float, int, Optional[Dict[str, int]]] _Frame = Tuple[float, float, float, int, Optional[Dict[str, int]]]
+67 -22
View File
@@ -28,6 +28,30 @@ import numpy as np
# long over one frame, so a sample this large is an idle gap between scrolls. # long over one frame, so a sample this large is an idle gap between scrolls.
FPS_LOG_INTERVAL = 5.0 FPS_LOG_INTERVAL = 5.0
# The stats line goes to INFO only when a window is worth an operator's
# attention, as Vegas's FPS line does (src/vegas_mode/coordinator.py): every
# 5s from every scroller was most of the journal on a healthy rig. A window is
# degraded when its frame rate falls below this fraction of the rate it was
# locked to (1 / its own median frame time; same 0.9 as Vegas) ...
STATS_HEALTHY_FRACTION = 0.9
# ... or when more than this share of its frames stalled (past 1.5x the
# median). A 1% stall rate barely moves the mean, so the fps test alone would
# miss the judder this line exists to show.
STATS_DEGRADED_STALL_RATE = 0.01
# A healthy scroller still logs at INFO this often, so silence in the journal
# means stopped rather than fine. Every window is still logged at DEBUG.
STATS_HEARTBEAT_INTERVAL = 300.0
def frame_stats_degraded(stats: Dict[str, Any]) -> bool:
"""Whether one frame_stats() window is worth logging at INFO."""
n = stats["frames"]
if n == 0 or stats["median"] <= 0:
return False
locked_fps = 1.0 / stats["median"]
return (stats["fps"] < locked_fps * STATS_HEALTHY_FRACTION
or stats["stalls"] > n * STATS_DEGRADED_STALL_RATE)
def _rgb_pixels(item) -> np.ndarray: def _rgb_pixels(item) -> np.ndarray:
"""An appended item's pixels as an RGB array, as pasting it would draw them.""" """An appended item's pixels as an RGB array, as pasting it would draw them."""
@@ -189,6 +213,11 @@ class ScrollHelper:
# Every frame time since the last stats line, so the 5s summary can # Every frame time since the last stats line, so the 5s summary can
# report the tail rather than one arbitrary sample. Cleared on log. # report the tail rather than one arbitrary sample. Cleared on log.
self._window: list = [] self._window: list = []
# INFO-level stats bookkeeping (see STATS_HEARTBEAT_INTERVAL). Kept
# across reset_scroll(): a heartbeat per scroll start would bring the
# chatter back. 0.0 so the first window after start-up is at INFO.
self._stats_last_info_log = 0.0
self._stats_was_degraded = False
# Scrolling state management # Scrolling state management
self.is_scrolling = False self.is_scrolling = False
@@ -532,7 +561,7 @@ class ScrollHelper:
width = self.display_width width = self.display_width
strip_width = self.cached_array.shape[1] strip_width = self.cached_array.shape[1]
if start_x + width + 1 <= strip_width: if 0 <= start_x and start_x + width + 1 <= strip_width:
# Slice the backing array directly. Going via # Slice the backing array directly. Going via
# _get_visible_portion_integer would build two PIL images only for # _get_visible_portion_integer would build two PIL images only for
# them to be converted straight back to arrays, which measured 15x # them to be converted straight back to arrays, which measured 15x
@@ -540,9 +569,10 @@ class ScrollHelper:
near = self.cached_array[:, start_x:start_x + width] near = self.cached_array[:, start_x:start_x + width]
far = self.cached_array[:, start_x + 1:start_x + 1 + width] far = self.cached_array[:, start_x + 1:start_x + 1 + width]
else: else:
# Close enough to the end that one of the slices wraps; let the # One of the slices wraps (close to the end, or a strip narrower
# integer path handle that and pay the conversion. Continuous mode # than the panel); let the integer path handle that and pay the
# extends the strip before reaching here, so this is the rare case. # conversion. Continuous mode extends the strip before reaching
# here, so this is the rare case.
near = np.asarray( near = np.asarray(
self._get_visible_portion_integer(start_x, start_x + width)) self._get_visible_portion_integer(start_x, start_x + width))
far = np.asarray( far = np.asarray(
@@ -572,29 +602,31 @@ class ScrollHelper:
_size = (self.display_width, self.display_height) _size = (self.display_width, self.display_height)
img_w = self.cached_array.shape[1] img_w = self.cached_array.shape[1]
if end_x <= img_w: if 0 <= start_x and end_x <= img_w:
# Normal case: single contiguous slice (fastest path) # Normal case: single contiguous slice (fastest path). tobytes()
frame_array = np.ascontiguousarray(self.cached_array[:, start_x:end_x]) # on the column-slice view already returns C-order bytes, so
return Image.frombytes('RGB', _size, frame_array.tobytes()) # ascontiguousarray() first only added a second full-frame copy.
else: return Image.frombytes(
'RGB', _size,
self.cached_array[:, start_x:end_x].tobytes())
# Ensure frame buffer is allocated for all non-simple paths # Ensure frame buffer is allocated for all non-simple paths
if self._frame_buffer is None or self._frame_buffer.shape != (self.display_height, self.display_width, 3): if self._frame_buffer is None or self._frame_buffer.shape != (self.display_height, self.display_width, 3):
self._frame_buffer = np.zeros((self.display_height, self.display_width, 3), dtype=np.uint8) self._frame_buffer = np.zeros((self.display_height, self.display_width, 3), dtype=np.uint8)
width1 = img_w - start_x if img_w == 0:
if width1 > 0: self._frame_buffer[:] = 0
# Wrap-around: tail of image + head of image
self._frame_buffer[:, :width1] = self.cached_array[:, start_x:]
remaining_width = self.display_width - width1
self._frame_buffer[:, width1:] = self.cached_array[:, :remaining_width]
else: else:
# Edge case: start_x at or past image end — show from beginning, # The frame runs off the strip, so it carries on from the head:
# clamped to available width (scroll_position should wrap before # frame column j is strip column (start_x + j) modulo the strip's
# reaching this state in normal operation). # width -- the tail and then the head, and a strip narrower than
available = min(self.display_width, img_w) # the panel repeated across it. Copying the tail and then the rest
self._frame_buffer[:, :available] = self.cached_array[:, :available] # of the frame from the head assumed the head was that wide, and
if available < self.display_width: # raised at every position for a strip narrower than the panel
self._frame_buffer[:, available:] = 0 # (Vegas composes one, with no lead-in, when its content is
# narrower than the chain).
np.take(self.cached_array, np.arange(start_x, end_x), axis=1,
mode='wrap', out=self._frame_buffer)
return Image.frombytes('RGB', _size, self._frame_buffer.tobytes()) return Image.frombytes('RGB', _size, self._frame_buffer.tobytes())
@@ -1207,10 +1239,23 @@ class ScrollHelper:
# as an idle gap. There is nothing to report, and reporting the # as an idle gap. There is nothing to report, and reporting the
# gap itself is the bug above. # gap itself is the bug above.
if self._window: if self._window:
# INFO when degraded, on the window that recovers from it, and
# as a slow heartbeat; DEBUG otherwise.
degraded = frame_stats_degraded(frame_stats(self._window))
if (degraded or self._stats_was_degraded
or current_time - self._stats_last_info_log
>= STATS_HEARTBEAT_INTERVAL):
self.logger.info( self.logger.info(
"Scroll frame stats - %s", "Scroll frame stats - %s",
format_frame_stats(self._window), format_frame_stats(self._window),
) )
self._stats_last_info_log = current_time
elif self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug(
"Scroll frame stats - %s",
format_frame_stats(self._window),
)
self._stats_was_degraded = degraded
self.last_fps_log_time = current_time self.last_fps_log_time = current_time
self.frame_count = 0 self.frame_count = 0
self._window = [] self._window = []
+43 -4
View File
@@ -18,7 +18,7 @@ the extra guard only stops a None size raising TypeError.
""" """
import logging import logging
from datetime import datetime, timezone from datetime import datetime, timedelta, timezone
from typing import Any, Dict, Optional, Tuple from typing import Any, Dict, Optional, Tuple
from zoneinfo import ZoneInfo from zoneinfo import ZoneInfo
@@ -338,10 +338,46 @@ def format_game_date(config: Optional[Dict[str, Any]], logger, date_text: str,
if not raw: if not raw:
return "" return ""
fmt = str(scroll_card_option(config, "date_format", "abbrev") or "abbrev") fmt = str(scroll_card_option(config, "date_format", "abbrev") or "abbrev")
return _format_date_as(fmt, raw, lambda: weekday_for(config, logger, game)) return _format_date_as(fmt, raw, lambda: weekday_for(config, logger, game),
game=game)
def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR) -> str: def _printed_weekday(game: Optional[Dict], month: int, day: int) -> str:
"""The weekday of the date a card prints as month/day, or '' if unknown.
The extractor prints "M/D" in the plugin's resolved zone (its own setting,
else the global one, else the system zone). The card cannot see that zone:
it is handed the plugin's config, whose ``timezone`` ships as "", so
card_tzinfo answers UTC and an evening kickoff in the Americas got the
next day's weekday ("Sat Oct 2" for a Friday game). Every zone is within
a day of UTC, so the printed date is the start's UTC date or a neighbour
of it; the one with that month and day is the date on the card.
"""
if not isinstance(game, dict):
return ""
raw = game.get("start_time_utc") or game.get("start_time")
if not raw:
return ""
try:
start = raw if isinstance(raw, datetime) else datetime.fromisoformat(
str(raw).replace("Z", "+00:00"))
if start.utcoffset() is None:
return "" # naive: no instant to place the date against
utc_day = start.astimezone(timezone.utc).date()
except (ValueError, TypeError, OverflowError):
return ""
for offset in (0, -1, 1):
try:
candidate = utc_day + timedelta(days=offset)
except OverflowError:
continue
if (candidate.month, candidate.day) == (month, day):
return WEEKDAY_ABBR[candidate.weekday()]
return ""
def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR,
game: Optional[Dict] = None) -> str:
"""Render a stripped, non-empty "M/D" *raw* in style *fmt*. """Render a stripped, non-empty "M/D" *raw* in style *fmt*.
The body both date formatters share. They differ in which setting names the The body both date formatters share. They differ in which setting names the
@@ -349,6 +385,9 @@ def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR) -> str:
``SportsCoreSharedMixin._format_game_date``), so those arrive as arguments: ``SportsCoreSharedMixin._format_game_date``), so those arrive as arguments:
*weekday* is a zero-argument callable, only called for the "weekday" style. *weekday* is a zero-argument callable, only called for the "weekday" style.
*months* lets the mixin keep reading its (overridable) ``_MONTH_ABBR``. *months* lets the mixin keep reading its (overridable) ``_MONTH_ABBR``.
With *game*, the "weekday" style names the printed date's own weekday
(:func:`_printed_weekday`), and *weekday* is only the fallback for a
date its start time cannot place.
""" """
if fmt == "numeric": if fmt == "numeric":
return raw return raw
@@ -364,7 +403,7 @@ def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR) -> str:
if fmt == "day_first": if fmt == "day_first":
return f"{day} {name}" return f"{day} {name}"
if fmt == "weekday": if fmt == "weekday":
day_name = weekday() day_name = _printed_weekday(game, month, day) or weekday()
return f"{day_name} {name} {day}" if day_name else f"{name} {day}" return f"{day_name} {name} {day}" if day_name else f"{name} {day}"
return f"{name} {day}" return f"{name} {day}"
+16 -2
View File
@@ -323,7 +323,7 @@ class SportsScrollDisplay:
:returns: True if a frame was drawn; False when there is no content or :returns: True if a frame was drawn; False when there is no content or
the frame could not be rendered. the frame could not be rendered.
""" """
if not self.scroll_helper.cached_image: if not self._has_strip():
return False return False
try: try:
@@ -416,7 +416,21 @@ class SportsScrollDisplay:
def has_cached_content(self) -> bool: def has_cached_content(self) -> bool:
"""Whether content is prepared and ready to scroll.""" """Whether content is prepared and ready to scroll."""
return bool(self.scroll_helper.cached_image) return self._has_strip()
def _has_strip(self) -> bool:
"""Whether the helper holds a strip, without building its PIL image.
Reading ``cached_image`` after the strip was extended or trimmed builds
the image from the array and keeps it, so the strip is held twice;
display_scroll_frame asks this every frame. ``has_strip()`` answers
from the helper's bookkeeping. A helper without it (a plugin's own, a
test double) is asked the old way.
"""
helper = self.scroll_helper
if callable(getattr(type(helper), "has_strip", None)):
return bool(helper.has_strip())
return bool(helper.cached_image)
# ------------------------------------------------------------------ # ------------------------------------------------------------------
# Live Vegas cards # Live Vegas cards
+4 -2
View File
@@ -360,14 +360,16 @@ class SportsCoreSharedMixin:
The formatting is sports_card's. What differs from the card's The formatting is sports_card's. What differs from the card's
``format_game_date`` is passed in: the setting (``switch_date_format``, ``format_game_date`` is passed in: the setting (``switch_date_format``,
see :meth:`_switch_date_format`) and the weekday, which comes from see :meth:`_switch_date_format`) and the weekday, which comes from
:meth:`_weekday_for` and so from this plugin's resolved timezone. :meth:`_weekday_for` and so from this plugin's resolved timezone
when the game's start cannot place the printed date. The game goes
in too, so both formatters name the printed date's own weekday.
""" """
raw = str(date_text or "").strip() raw = str(date_text or "").strip()
if not raw: if not raw:
return raw return raw
return _card._format_date_as(self._switch_date_format(), raw, return _card._format_date_as(self._switch_date_format(), raw,
lambda: self._weekday_for(game), lambda: self._weekday_for(game),
self._MONTH_ABBR) self._MONTH_ABBR, game=game)
def _weekday_for(self, game: Optional[Dict]) -> str: def _weekday_for(self, game: Optional[Dict]) -> str:
"""Weekday abbreviation from the game's start time, or ''.""" """Weekday abbreviation from the game's start time, or ''."""
+6 -1
View File
@@ -29,7 +29,6 @@ import time
import logging import logging
from enum import Enum from enum import Enum
from typing import Callable, Optional from typing import Callable, Optional
import numpy as np
from PIL import Image from PIL import Image
from src.config_manager_atomic import _replace from src.config_manager_atomic import _replace
@@ -434,6 +433,12 @@ class DisplaySyncManager:
return return
if self._leader_state != LeaderState.CONNECTED or not self._peer_ip: if self._leader_state != LeaderState.CONNECTED or not self._peer_ip:
return return
# numpy is imported here, not at module level: the web interface
# imports this module for its constants (STATUS_FILE, SYNC_PORT) and
# would otherwise load numpy for nothing. Only a connected leader
# gets this far, and after the first frame the import is a
# sys.modules lookup.
import numpy as np
try: try:
arr = np.asarray(image.convert("RGB"), dtype=np.uint8) arr = np.asarray(image.convert("RGB"), dtype=np.uint8)
header = _RAW_MAGIC + _RAW_HEADER.pack(image.width, image.height) header = _RAW_MAGIC + _RAW_HEADER.pack(image.width, image.height)
+61 -11
View File
@@ -14,7 +14,7 @@ import json
import time import time
import threading import threading
from pathlib import Path from pathlib import Path
from typing import Dict, Any, Optional, List, Callable from typing import Dict, Any, Optional, List, Callable, Tuple
from collections import defaultdict from collections import defaultdict
import logging import logging
import hashlib import hashlib
@@ -52,6 +52,17 @@ class ConfigService:
# Thread safety # Thread safety
self._lock: threading.RLock = threading.RLock() self._lock: threading.RLock = threading.RLock()
# Held across a whole reload -- read, swap, notify -- so one reload's
# notifications finish before the next one's start. Subscribers run
# under this lock and never under _lock: the display's per-plugin
# subscriber can wait seconds for a busy plugin, and get_config(),
# subscribe() and unsubscribe() -- called from the render thread --
# must not wait behind it.
self._notify_lock: threading.RLock = threading.RLock()
# (key, callback, thread id) of the callback a notification is running,
# so unsubscribe() can wait for that one call; signalled on its return.
self._running_callback: Optional[Tuple[str, Callable[..., None], int]] = None
self._callback_done = threading.Condition(self._lock)
# Current configuration # Current configuration
self._current_config: Dict[str, Any] = {} self._current_config: Dict[str, Any] = {}
@@ -87,6 +98,7 @@ class ConfigService:
True if config changed, False otherwise True if config changed, False otherwise
""" """
try: try:
with self._notify_lock:
new_config = self.config_manager.load_config() new_config = self.config_manager.load_config()
new_checksum = self._calculate_checksum(new_config) new_checksum = self._calculate_checksum(new_config)
@@ -103,7 +115,7 @@ class ConfigService:
self._current_config = new_config self._current_config = new_config
self._current_checksum = new_checksum self._current_checksum = new_checksum
# Notify subscribers # Notify subscribers, outside _lock (see _notify_lock)
self._notify_subscribers(old_config, new_config) self._notify_subscribers(old_config, new_config)
self.logger.info( self.logger.info(
@@ -127,16 +139,19 @@ class ConfigService:
Args: Args:
old_config: Previous configuration old_config: Previous configuration
new_config: New configuration new_config: New configuration
Called without _lock held. The subscriber lists are copied under it,
and each callback is checked against them again just before it runs.
""" """
with self._lock:
subscribers = {key: list(callbacks) for key, callbacks in self._subscribers.items()}
# Notify global subscribers (key: '*') # Notify global subscribers (key: '*')
for callback in self._subscribers.get('*', []): for callback in subscribers.get('*', []):
try: self._call_subscriber('*', callback, old_config, new_config)
callback(old_config, new_config)
except Exception as e:
self.logger.error("Error in global config change callback: %s", e, exc_info=True)
# Notify plugin-specific subscribers # Notify plugin-specific subscribers
for plugin_id in self._subscribers.keys(): for plugin_id, callbacks in subscribers.items():
if plugin_id == '*': if plugin_id == '*':
continue continue
@@ -145,16 +160,42 @@ class ConfigService:
# Only notify if plugin config actually changed # Only notify if plugin config actually changed
if old_plugin_config != new_plugin_config: if old_plugin_config != new_plugin_config:
for callback in self._subscribers[plugin_id]: for callback in callbacks:
self._call_subscriber(plugin_id, callback,
old_plugin_config, new_plugin_config)
def _call_subscriber(
self,
key: str,
callback: Callable[[Dict[str, Any], Dict[str, Any]], None],
old_config: Dict[str, Any],
new_config: Dict[str, Any],
) -> None:
"""Run one callback, unless it was unsubscribed since the snapshot.
unsubscribe() promises that once it returns the callback is neither
running nor will run: the display unloads the plugin straight after.
"""
with self._lock:
if callback not in self._subscribers.get(key, ()):
return
self._running_callback = (key, callback, threading.get_ident())
try: try:
callback(old_plugin_config, new_plugin_config) callback(old_config, new_config)
except Exception as e: except Exception as e:
if key == '*':
self.logger.error("Error in global config change callback: %s", e, exc_info=True)
else:
self.logger.error( self.logger.error(
"Error in config change callback for %s: %s", "Error in config change callback for %s: %s",
plugin_id, key,
e, e,
exc_info=True exc_info=True
) )
finally:
with self._lock:
self._running_callback = None
self._callback_done.notify_all()
def _check_file_changes(self) -> bool: def _check_file_changes(self) -> bool:
""" """
@@ -276,6 +317,11 @@ class ConfigService:
""" """
Unsubscribe from configuration changes. Unsubscribe from configuration changes.
Once this returns the callback is not running and will not be called
again. A notification that is running this very callback is waited
for (unless the callback is the caller); one running any other
callback is not.
Args: Args:
callback: Callback function to remove callback: Callback function to remove
plugin_id: Optional plugin ID (must match subscription) plugin_id: Optional plugin ID (must match subscription)
@@ -285,6 +331,10 @@ class ConfigService:
if callback in self._subscribers[key]: if callback in self._subscribers[key]:
self._subscribers[key].remove(callback) self._subscribers[key].remove(callback)
self.logger.debug("Unsubscribed from config changes for %s", key) self.logger.debug("Unsubscribed from config changes for %s", key)
while (self._running_callback is not None
and self._running_callback[:2] == (key, callback)
and self._running_callback[2] != threading.get_ident()):
self._callback_done.wait()
def shutdown(self) -> None: def shutdown(self) -> None:
"""Shutdown the configuration service.""" """Shutdown the configuration service."""
+182 -16
View File
@@ -25,11 +25,12 @@ import os
import inspect import inspect
import signal import signal
import json import json
import math
import threading import threading
import types import types
from collections import deque from collections import deque
from contextlib import contextmanager from contextlib import contextmanager
from typing import Dict, Any, List, Optional, Callable, Set, Tuple from typing import Dict, Any, FrozenSet, List, Optional, Callable, Set, Tuple
from datetime import datetime from datetime import datetime
from concurrent.futures import ThreadPoolExecutor, as_completed # pylint: disable=no-name-in-module from concurrent.futures import ThreadPoolExecutor, as_completed # pylint: disable=no-name-in-module
import pytz import pytz
@@ -55,7 +56,7 @@ from src.ipc.contract import (
PluginReloadArgs, PluginReloadArgs,
PluginReloadResult, PluginReloadResult,
) )
from src.ipc.server import ControlServer, QueuedCommand, start_control_server from src.ipc.server import ControlServer, QueuedCommand, StateHub, start_control_server
from src.vegas_mode.render_pipeline import SYNC_SEND_INTERVAL from src.vegas_mode.render_pipeline import SYNC_SEND_INTERVAL
# Get logger with consistent configuration # Get logger with consistent configuration
@@ -65,6 +66,13 @@ logger = get_logger(__name__)
# treats display_current_state older than 120 s as unknown. # treats display_current_state older than 120 s as unknown.
CURRENT_STATE_REFRESH_SECONDS = 30 CURRENT_STATE_REFRESH_SECONDS = 30
# While the control socket serves the web interface's state readers
# (StateHub.readers_active), display_current_state is only their fallback:
# it is then rewritten at this interval and on a change of the flags, not on
# every mode change. Below the readers' 120 s max_age, so the fallback copy
# never reads as unknown.
CURRENT_STATE_RELAXED_REFRESH_SECONDS = 60
# How long startup will wait for plugins to fetch their first data before # How long startup will wait for plugins to fetch their first data before
# showing anything. Each plugin's update blocks for up to the executor's 30s # showing anything. Each plugin's update blocks for up to the executor's 30s
# timeout and they run one after another, so the uncapped total is the sum of # timeout and they run one after another, so the uncapped total is the sum of
@@ -82,6 +90,19 @@ _MIN_INITIAL_UPDATE_TIMEOUT_SECONDS = 2.0
DEFAULT_DYNAMIC_DURATION_CAP = 180.0 DEFAULT_DYNAMIC_DURATION_CAP = 180.0
def _finite_seconds(value: Any) -> Optional[float]:
"""``value`` as seconds when it is a finite number or a numeric string,
else None. A bool is not a number here, though it is an int: True would
read as a one-second screen."""
if isinstance(value, bool):
return None
try:
seconds = float(value)
except (TypeError, ValueError, OverflowError):
return None
return seconds if math.isfinite(seconds) else None
class _PluginReloadJob: class _PluginReloadJob:
"""A ``plugin.reload`` whose slow half runs off the render thread. """A ``plugin.reload`` whose slow half runs off the render thread.
@@ -323,6 +344,9 @@ class DisplayController:
# Monotonic stamp of the last _service_pending_changes pass; same # Monotonic stamp of the last _service_pending_changes pass; same
# "None means never" convention as _last_on_demand_poll. # "None means never" convention as _last_on_demand_poll.
self._last_pending_service: Optional[float] = None self._last_pending_service: Optional[float] = None
# Monotonic stamp of the last scheduled-update pass; see
# _tick_plugin_updates_if_due. Same "None means never" convention.
self._last_plugin_update_tick: Optional[float] = None
# The control socket (src/ipc), started by run(). None when it is not # The control socket (src/ipc), started by run(). None when it is not
# served (Windows, LEDMATRIX_CONTROL_SOCKET=off, a bind failure); # served (Windows, LEDMATRIX_CONTROL_SOCKET=off, a bind failure);
# the file mailbox works either way. # the file mailbox works either way.
@@ -1089,10 +1113,37 @@ class DisplayController:
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Error marking plugin %s updated for Vegas", plugin_id) logger.exception("Error marking plugin %s updated for Vegas", plugin_id)
#: Shortest gap between scheduled-update passes from the frame loops and
#: the dwell sleep. The pass (PluginManager.run_scheduled_updates) copies
#: the plugin dict and takes several locks per plugin to find, almost
#: always, that nothing is due: about 95 us with 20 plugins on a Pi, or
#: 1.2% of the render thread at 125 frames a second. No interval is
#: shorter than PluginManager.MIN_DYNAMIC_UPDATE_INTERVAL (5 s), and the
#: 1 Hz frame loop already ticks once a second, so a quarter second late
#: is not noticed.
PLUGIN_UPDATE_TICK_INTERVAL = 0.25
#: Class-level default for controllers built without __init__ (tests).
_last_plugin_update_tick: Optional[float] = None
def _tick_plugin_updates_if_due(self) -> None:
"""_tick_plugin_updates, at most once per PLUGIN_UPDATE_TICK_INTERVAL.
For the per-frame callers. The top of each loop pass calls
_tick_plugin_updates itself, unthrottled, because that is where a
plugin just loaded, reloaded or enabled for on-demand gets its first
update, and it must not wait out the floor.
"""
last = self._last_plugin_update_tick
if last is not None and time.monotonic() - last < self.PLUGIN_UPDATE_TICK_INTERVAL:
return
self._tick_plugin_updates()
def _tick_plugin_updates(self): def _tick_plugin_updates(self):
"""Run any plugin updates that are due.""" """Run any plugin updates that are due."""
if not self.plugin_manager: if not self.plugin_manager:
return return
self._last_plugin_update_tick = time.monotonic()
try: try:
self.plugin_manager.run_scheduled_updates() self.plugin_manager.run_scheduled_updates()
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
@@ -1260,7 +1311,7 @@ class DisplayController:
# A dwell can be a minute long (sixty seconds while scheduled # A dwell can be a minute long (sixty seconds while scheduled
# off); the watchdog must hear from this thread throughout. # off); the watchdog must hear from this thread throughout.
display_watchdog.watchdog.beat() display_watchdog.watchdog.beat()
self._tick_plugin_updates() self._tick_plugin_updates_if_due()
self._service_pending_changes() self._service_pending_changes()
self._check_live_takeover() self._check_live_takeover()
if (self.current_display_mode != mode if (self.current_display_mode != mode
@@ -1303,6 +1354,12 @@ class DisplayController:
"until one does", self.EMPTY_ROTATION_PAUSE) "until one does", self.EMPTY_ROTATION_PAUSE)
self._sleep_with_plugin_updates(self.EMPTY_ROTATION_PAUSE) self._sleep_with_plugin_updates(self.EMPTY_ROTATION_PAUSE)
#: Plugins already warned about a display duration that is not a number,
#: so a bad setting logs once, not at every one of its screens. A
#: frozenset, replaced rather than mutated; class-level default for
#: controllers built without __init__ (tests).
_duration_warned: FrozenSet[str] = frozenset()
def _get_display_duration(self, mode_key): def _get_display_duration(self, mode_key):
"""Seconds to show a mode: the Rotation & Durations page's value for it """Seconds to show a mode: the Rotation & Durations page's value for it
(display.display_durations), else the plugin's own duration. (display.display_durations), else the plugin's own duration.
@@ -1310,6 +1367,17 @@ class DisplayController:
The saved value has to win. Every plugin inherits The saved value has to win. Every plugin inherits
get_display_duration(), so checking the plugin first meant the page's get_display_duration(), so checking the plugin first meant the page's
values were never read. values were never read.
The plugin's answer is checked here, not trusted. Several plugins
return their display_duration setting straight from config.json, so
one saved as "20" or null (the raw config editor, a hand edit) came
back as a string or None; _resolve_durations compared it with 0, and
the TypeError went past every handler in the loop and stopped the
display service, which systemd restarted into the same screen. A
numeric string counts, as in BasePlugin.get_display_duration; any
other value that is not a finite number, or a raise, gets the 30 s a
mode without a plugin gets. A number at or below zero is passed on:
_resolve_durations has its own rule for that.
""" """
display_durations = self.config.get('display', {}).get('display_durations', {}) or {} display_durations = self.config.get('display', {}).get('display_durations', {}) or {}
override = display_durations.get(mode_key) override = display_durations.get(mode_key)
@@ -1317,8 +1385,22 @@ class DisplayController:
return float(override) return float(override)
plugin_instance = self.plugin_modes.get(mode_key) plugin_instance = self.plugin_modes.get(mode_key)
if plugin_instance is not None and hasattr(plugin_instance, 'get_display_duration'): if plugin_instance is None or not hasattr(plugin_instance, 'get_display_duration'):
return plugin_instance.get_display_duration() return 30
try:
value = plugin_instance.get_display_duration()
except Exception as err: # pylint: disable=broad-except
problem = f"get_display_duration() raised {type(err).__name__}: {err}"
else:
seconds = _finite_seconds(value)
if seconds is not None:
return seconds
problem = f"display duration {value!r} is not a number"
plugin_id = getattr(plugin_instance, 'plugin_id', None) or mode_key
if plugin_id not in self._duration_warned:
self._duration_warned = self._duration_warned | {plugin_id}
logger.warning("Plugin %s: %s; showing its modes for 30s (logged once)",
plugin_id, problem)
return 30 return 30
def _get_global_dynamic_cap(self) -> Optional[float]: def _get_global_dynamic_cap(self) -> Optional[float]:
@@ -1430,10 +1512,14 @@ class DisplayController:
return None return None
return max(0.0, expires_at - time.time()) return max(0.0, expires_at - time.time())
def _publish_current_mode_state(self) -> None: #: The control socket's state stream (src/ipc/server.StateHub), while the
"""Publish the currently active display mode/plugin to cache for the web UI.""" #: socket is served. Class-level default for controllers built without
try: #: __init__ (tests) and for a display with no socket.
state = { _state_hub: Optional[StateHub] = None
def _current_mode_state(self) -> Dict[str, Any]:
"""What display_current_state and the socket's ``display`` section hold."""
return {
'mode': self.current_display_mode, 'mode': self.current_display_mode,
'plugin_id': self.mode_to_plugin_id.get(self.current_display_mode), 'plugin_id': self.mode_to_plugin_id.get(self.current_display_mode),
'mode_index': self.current_mode_index, 'mode_index': self.current_mode_index,
@@ -1442,6 +1528,44 @@ class DisplayController:
'is_display_active': self.is_display_active, 'is_display_active': self.is_display_active,
'last_updated': time.time(), 'last_updated': time.time(),
} }
def _push_live_state(self, display_state: Optional[Dict[str, Any]] = None) -> None:
"""Hand the current mode and the brightness to the control socket's
state stream. In memory, no disk: the hub only bumps its version (and
wakes subscribers) when something other than ``last_updated`` changed.
Called on every pass of the publish points below, so
``display.last_updated`` doubles as the render thread's proof of life
for the socket's readers, as the cache key's max_age does today.
"""
hub = self._state_hub
if hub is None:
return
try:
hub.publish('display', display_state or self._current_mode_state(),
volatile=('last_updated',))
hub.publish('brightness', {
'brightness': getattr(self, '_normal_brightness', None),
'panel_brightness': getattr(self, 'current_brightness', None),
'dimmed': bool(getattr(self, 'is_dimmed', False)),
})
except Exception as err: # pylint: disable=broad-except
logger.debug("Could not publish the display state to the control socket: %s",
err, exc_info=True)
def _state_readers_on_socket(self) -> bool:
"""Is the control socket serving the web interface's state readers?"""
hub = self._state_hub
try:
return bool(hub is not None and hub.readers_active())
except Exception: # pylint: disable=broad-except
return False
def _publish_current_mode_state(self) -> None:
"""Publish the currently active display mode/plugin to cache for the web UI."""
try:
state = self._current_mode_state()
self._push_live_state(state)
self.cache_manager.set('display_current_state', state) self.cache_manager.set('display_current_state', state)
self._last_published_mode = self.current_display_mode self._last_published_mode = self.current_display_mode
self._last_published_flags = self._current_state_flags() self._last_published_flags = self._current_state_flags()
@@ -1467,11 +1591,23 @@ class DisplayController:
priority, a single enabled plugin -- has to be republished or the UI priority, a single enabled plugin -- has to be republished or the UI
reports it as unknown. Otherwise this writes only on a change, not on reports it as unknown. Otherwise this writes only on a change, not on
every render tick. every render tick.
While the control socket serves the web interface's state readers,
the socket's in-memory copy is updated on every call and the cache
key is only their fallback: a mode change alone is then written at
the relaxed refresh (CURRENT_STATE_RELAXED_REFRESH_SECONDS), still
inside the readers' max_age. The flags are still written at once.
When the socket stops serving them, the next call writes a changed
mode again.
""" """
if (self.current_display_mode != self._last_published_mode relaxed = self._state_readers_on_socket()
refresh = CURRENT_STATE_RELAXED_REFRESH_SECONDS if relaxed else CURRENT_STATE_REFRESH_SECONDS
if ((not relaxed and self.current_display_mode != self._last_published_mode)
or self._current_state_flags() != getattr(self, '_last_published_flags', None) or self._current_state_flags() != getattr(self, '_last_published_flags', None)
or time.monotonic() - self._last_published_at >= CURRENT_STATE_REFRESH_SECONDS): or time.monotonic() - self._last_published_at >= refresh):
self._publish_current_mode_state() self._publish_current_mode_state() # pushes to the socket as well
else:
self._push_live_state()
def _on_demand_state(self) -> Dict[str, Any]: def _on_demand_state(self) -> Dict[str, Any]:
"""The on-demand state as published to the cache and the control socket.""" """The on-demand state as published to the cache and the control socket."""
@@ -1494,6 +1630,11 @@ class DisplayController:
"""Publish current on-demand state to cache for external consumers.""" """Publish current on-demand state to cache for external consumers."""
try: try:
state = self._on_demand_state() state = self._on_demand_state()
hub = self._state_hub
if hub is not None:
# In memory, first: a subscriber hears the outcome of an
# on-demand command even if the cache write below fails.
hub.publish('on_demand', state, volatile=('last_updated', 'remaining'))
self.cache_manager.set('display_on_demand_state', state) self.cache_manager.set('display_on_demand_state', state)
except (OSError, RuntimeError, ValueError, TypeError) as err: except (OSError, RuntimeError, ValueError, TypeError) as err:
logger.error("Failed to publish on-demand state: %s", err, exc_info=True) logger.error("Failed to publish on-demand state: %s", err, exc_info=True)
@@ -1738,12 +1879,31 @@ class DisplayController:
""" """
if self._control_server is not None: if self._control_server is not None:
return return
hub = StateHub(loop_probe=display_watchdog.watchdog.liveness)
try: try:
self._control_server = start_control_server( self._control_server = start_control_server(
status_provider=self._control_status, status_provider=self._control_status,
cache_dir=getattr(self.cache_manager, 'cache_dir', None)) cache_dir=getattr(self.cache_manager, 'cache_dir', None),
state_hub=hub)
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Control socket not started; using the file mailbox only") logger.exception("Control socket not started; using the file mailbox only")
if self._control_server is not None:
self._start_state_stream(hub)
def _start_state_stream(self, hub: StateHub) -> None:
"""Start publishing to the socket's state stream (``state.get`` and
``state.subscribe``): everything a reader would see, now, then on
every publish. Never raises; without it readers use the cache keys."""
try:
self._state_hub = hub
self._push_live_state()
hub.publish('on_demand', self._on_demand_state(),
volatile=('last_updated', 'remaining'))
publisher = getattr(self, '_plugin_runtime_publisher', None)
if publisher is not None:
publisher.attach_hub(hub)
except Exception: # pylint: disable=broad-except
logger.exception("Control socket state stream not started; readers use the cache")
def _control_status(self) -> Dict[str, Any]: def _control_status(self) -> Dict[str, Any]:
"""The socket's on_demand.status answer. Runs on the socket's thread: reads only.""" """The socket's on_demand.status answer. Runs on the socket's thread: reads only."""
@@ -2820,7 +2980,10 @@ class DisplayController:
self._follower_local_x = local_x self._follower_local_x = local_x
if rp and rp.scroll_helper.cached_image is not None: # has_strip(), not cached_image: every follower frame asks, and
# reading a strip the follower's own rebuild deferred would build
# and keep a second copy of it as a PIL image.
if rp and rp.scroll_helper.has_strip():
# Hold last frame until TCP image arrives after cycle reset # Hold last frame until TCP image arrives after cycle reset
if not self._follower_pending_new_image and local_x >= width: if not self._follower_pending_new_image and local_x >= width:
rp.scroll_helper.scroll_position = ( rp.scroll_helper.scroll_position = (
@@ -3510,6 +3673,8 @@ class DisplayController:
self._release_on_demand_plugins() self._release_on_demand_plugins()
if not self.available_modes: if not self.available_modes:
continue # it was all there was; idle as above continue # it was all there was; idle as above
# Unthrottled, unlike the frame loops: a plugin loaded,
# reloaded or enabled for on-demand above is due at once.
self._tick_plugin_updates() self._tick_plugin_updates()
# Clean up expired WiFi status messages # Clean up expired WiFi status messages
@@ -3772,7 +3937,7 @@ class DisplayController:
# Multi-display sync: send follower frame after each render # Multi-display sync: send follower frame after each render
self._send_follower_frame(manager_to_display) self._send_follower_frame(manager_to_display)
self._tick_plugin_updates() self._tick_plugin_updates_if_due()
# Throttled: one clock compare between passes. # Throttled: one clock compare between passes.
self._service_pending_changes() self._service_pending_changes()
self._check_live_takeover() self._check_live_takeover()
@@ -3831,7 +3996,7 @@ class DisplayController:
"breaking early", active_mode, "breaking early", active_mode,
self.current_display_mode) self.current_display_mode)
break break
self._tick_plugin_updates() self._tick_plugin_updates_if_due()
elapsed = time.time() - start_time elapsed = time.time() - start_time
if elapsed >= target_duration: if elapsed >= target_duration:
@@ -4532,6 +4697,7 @@ class DisplayController:
except Exception as e: except Exception as e:
logger.warning("Error closing the control socket: %s", e) logger.warning("Error closing the control socket: %s", e)
self._control_server = None self._control_server = None
self._state_hub = None
# Stop the async update worker first so no in-flight update() call # Stop the async update worker first so no in-flight update() call
# is still touching display/cache-backed resources while they're # is still touching display/cache-backed resources while they're
# torn down below. # torn down below.
+4 -1
View File
@@ -57,7 +57,8 @@ import freetype
from src.common import snapshot_policy from src.common import snapshot_policy
from src import display_watchdog from src import display_watchdog
from src.common.frame_timing import FrameTimingRecorder, install_gc_monitor from src.common.frame_timing import (
FrameTimingRecorder, install_gc_monitor, uninstall_gc_monitor)
if TYPE_CHECKING: if TYPE_CHECKING:
from src.common.render_gate import RenderGate from src.common.render_gate import RenderGate
@@ -1407,6 +1408,8 @@ class DisplayManager:
# The stall watchdog would otherwise outlive this manager. # The stall watchdog would otherwise outlive this manager.
if getattr(self, 'frame_timing', None) is not None: if getattr(self, 'frame_timing', None) is not None:
self.frame_timing.close() self.frame_timing.close()
# Installed with the recorder; stop timing collections with it.
uninstall_gc_monitor()
# Reset the singleton state when cleaning up # Reset the singleton state when cleaning up
DisplayManager._instance = None DisplayManager._instance = None
+17
View File
@@ -216,6 +216,23 @@ class RenderWatchdog:
def armed(self) -> bool: def armed(self) -> bool:
return self._armed return self._armed
def liveness(self) -> Dict[str, Any]:
"""The heartbeat, in memory: what the control socket's state stream
reports as ``loop``.
``heartbeat_age_seconds`` is the age of the render thread's last beat,
the beat that writes the heartbeat file, so it ages at the same rate
and is judged by the same ``HEARTBEAT_STALE_SECONDS``. None until the
loop has drawn its first frame. Any thread may call this: it only
reads two attributes.
"""
last = self._last_beat
age = None
if self._armed and last is not None:
age = max(self._clock() - last, 0.0)
return {'heartbeat_age_seconds': age, 'armed': self._armed,
'stale_after': HEARTBEAT_STALE_SECONDS}
def _on_render_thread(self) -> bool: def _on_render_thread(self) -> bool:
return self._render_thread is not None and threading.get_ident() == self._render_thread return self._render_thread is not None and threading.get_ident() == self._render_thread
+2 -1
View File
@@ -23,6 +23,7 @@ import requests
from typing import Any, Dict, List from typing import Any, Dict, List
from src.common.api_helper import DEFAULT_HTTP_HEADERS from src.common.api_helper import DEFAULT_HTTP_HEADERS
from src.common.json_body import response_json
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -157,7 +158,7 @@ class DynamicTeamResolver:
response = requests.get(rankings_url, headers=dict(DEFAULT_HTTP_HEADERS), response = requests.get(rankings_url, headers=dict(DEFAULT_HTTP_HEADERS),
timeout=self.request_timeout) timeout=self.request_timeout)
response.raise_for_status() response.raise_for_status()
data = response.json() data = response_json(response)
rankings = {} rankings = {}
rankings_data = data.get('rankings', []) rankings_data = data.get('rankings', [])
+1 -1
View File
@@ -485,7 +485,7 @@ def record_error(
# and only the display service's ever records anything (plugin_executor runs # and only the display service's ever records anything (plugin_executor runs
# the plugins there). The web interface therefore reads a snapshot the display # the plugins there). The web interface therefore reads a snapshot the display
# service publishes to the shared cache directory -- the same channel, and the # service publishes to the shared cache directory -- the same channel, and the
# same file permissions, as display_current_state and plugin_metrics:*: files # same file permissions, as display_current_state and plugin_metrics_snapshot: files
# are 0660 and carry the cache directory's group, so root writes and the web # are 0660 and carry the cache directory's group, so root writes and the web
# user reads, and the other way round for the clear request. # user reads, and the other way round for the clear request.
# #
+262 -1
View File
@@ -10,20 +10,24 @@ blocks for longer than ``timeout`` in total.
from __future__ import annotations from __future__ import annotations
import socket import socket
import threading
import time import time
import uuid import uuid
from typing import Any, Dict, List, Mapping, Optional, Sequence from typing import Any, Callable, Dict, List, Mapping, Optional, Sequence
from src.ipc.contract import ( from src.ipc.contract import (
AWAIT_SECONDS, AWAIT_SECONDS,
MAX_MESSAGE_BYTES, MAX_MESSAGE_BYTES,
PROTOCOL_VERSION, PROTOCOL_VERSION,
SUBSCRIBE_KEEPALIVE_SECONDS,
SUPPORTED_VERSIONS, SUPPORTED_VERSIONS,
Command, Command,
FrameReader, FrameReader,
ProtocolError, ProtocolError,
Request, Request,
Response, Response,
StateEvent,
StateEventKind,
client_socket_paths, client_socket_paths,
decode_message, decode_message,
encode_message, encode_message,
@@ -233,3 +237,260 @@ def hello(client: str = 'web', *, timeout: float = DEFAULT_TIMEOUT_SECONDS,
"""Version negotiation: the result's ``version`` is the one both sides speak.""" """Version negotiation: the result's ``version`` is the one both sides speak."""
return request(Command.HELLO, {'versions': list(SUPPORTED_VERSIONS), 'client': client}, return request(Command.HELLO, {'versions': list(SUPPORTED_VERSIONS), 'client': client},
timeout=timeout, paths=paths) timeout=timeout, paths=paths)
# -- the state stream (stage 3) ---------------------------------------------------------
def state_get(since: Optional[int] = None, epoch: Optional[str] = None, *,
timeout: float = DEFAULT_TIMEOUT_SECONDS,
paths: Optional[Sequence[str]] = None) -> Dict[str, Any]:
"""The display's state now, as a :class:`~src.ipc.contract.StateSnapshot`.
With ``since``/``epoch`` from an earlier answer, an unchanged state comes
back in the short ``changed: false`` form. Raises :class:`ControlError`
(``unknown_command`` from a display older than stage 3).
"""
args: Dict[str, Any] = {}
if since is not None:
args['since'] = since
if epoch is not None:
args['epoch'] = epoch
return request(Command.STATE_GET, args, timeout=timeout, paths=paths)
def snapshot_age(snapshot: Mapping[str, Any], now_mono: Optional[float] = None) -> float:
"""Seconds since ``snapshot`` arrived: ``received_mono`` (set by
:meth:`StateSubscription.latest`) to now; 0 for a one-shot answer."""
received = snapshot.get('received_mono')
if isinstance(received, (int, float)) and not isinstance(received, bool):
now_mono = time.monotonic() if now_mono is None else now_mono
return max(now_mono - float(received), 0.0)
return 0.0
def snapshot_loop_age(snapshot: Mapping[str, Any],
now_mono: Optional[float] = None) -> Optional[float]:
"""The render loop's heartbeat age now, from a state snapshot: the age the
display measured when it answered, plus the time since the answer
arrived. None when the display has no beat to report yet."""
loop = snapshot.get('loop')
if not isinstance(loop, dict):
state = snapshot.get('state')
loop = state.get('loop') if isinstance(state, dict) else None
age = loop.get('heartbeat_age_seconds') if isinstance(loop, dict) else None
if not isinstance(age, (int, float)) or isinstance(age, bool):
return None
return max(float(age), 0.0) + snapshot_age(snapshot, now_mono)
def _merge_volatile(state: Dict[str, Any], volatile: Any) -> None:
"""Fold a tick's ``volatile`` values (``{section: {key: value}}``) into
``state``, copying each section it touches.
These are the timestamps the hub leaves out of its version --
``display.last_updated``, ``on_demand.last_updated``/``remaining``,
``plugins.published_at`` -- and the readers judge freshness by them, so
a copy that only full ``state`` events updated would go stale while the
same mode stayed on screen. Only keys the section already has are taken:
a tick never adds a section or a key the last snapshot did not carry
(a section left out of a truncated snapshot stays out).
"""
if not isinstance(volatile, dict):
return # a display from before ticks carried them
for name, values in volatile.items():
section = state.get(name)
if not isinstance(section, dict) or not isinstance(values, dict):
continue
fresh = {k: v for k, v in values.items() if k in section}
if fresh:
state[name] = dict(section, **fresh)
#: A subscription that has heard nothing for this long is not trusted: the
#: display sends a tick at least every SUBSCRIBE_KEEPALIVE_SECONDS.
SUBSCRIPTION_SILENCE_SECONDS = 3 * SUBSCRIBE_KEEPALIVE_SECONDS
#: Reconnect backoff: the first retry, and the cap. A display that does not
#: know state.subscribe (stage 2 or older) is retried at the cap.
_RECONNECT_MIN_SECONDS = 1.0
_RECONNECT_MAX_SECONDS = 30.0
#: Failures that another try soon will not fix.
_SLOW_RETRY_REASONS = frozenset({'unknown_command', 'unsupported_version', 'disabled',
'unsupported'})
class StateSubscription:
"""One ``state.subscribe`` connection, held on a daemon thread.
Keeps the latest snapshot the display pushed, so a reader answers from
memory (:meth:`latest`). Reconnects with a backoff when the display goes
away. Never raises into the caller: :meth:`latest` is None whenever the
copy cannot be vouched for (not connected, or silent for longer than
``silence``), and the caller falls back.
"""
def __init__(self, paths: Optional[Sequence[str]] = None, *,
silence: float = SUBSCRIPTION_SILENCE_SECONDS,
connect_timeout: float = DEFAULT_TIMEOUT_SECONDS,
clock: Callable[[], float] = time.monotonic):
self._paths = list(paths) if paths is not None else None
self._silence = silence
self._connect_timeout = connect_timeout
self._clock = clock
self._lock = threading.Lock()
self._snapshot: Optional[Dict[str, Any]] = None
self._received: Optional[float] = None
self._connected = False
self._stop = threading.Event()
self._sock: Optional[socket.socket] = None
self._thread: Optional[threading.Thread] = None
#: The reason the last connection ended (a ControlError reason).
self.last_error: Optional[str] = None
#: Full snapshots received: the subscribe answer and each state event.
self.snapshots = 0
# -- the reader's side ---------------------------------------------------
@property
def connected(self) -> bool:
return self._connected
def latest(self) -> Optional[Dict[str, Any]]:
"""A copy of the latest snapshot, with ``received_mono`` (this
process's monotonic clock when it arrived); None when not trusted."""
with self._lock:
if not self._connected or self._snapshot is None or self._received is None:
return None
if self._clock() - self._received > self._silence:
return None
snap = dict(self._snapshot)
snap['received_mono'] = self._received
return snap
# -- lifecycle -----------------------------------------------------------
def start(self) -> 'StateSubscription':
if self._thread is None or not self._thread.is_alive():
self._stop.clear()
self._thread = threading.Thread(target=self._run, name='ledmatrix-state-feed',
daemon=True)
self._thread.start()
return self
def stop(self, timeout: float = 2.0) -> None:
self._stop.set()
sock = self._sock
if sock is not None:
try:
sock.shutdown(socket.SHUT_RDWR)
except OSError:
pass
thread = self._thread
if thread is not None and thread is not threading.current_thread():
thread.join(timeout)
self._thread = None
# -- the feed thread -----------------------------------------------------
def _run(self) -> None:
backoff = _RECONNECT_MIN_SECONDS
while not self._stop.is_set():
snapshots = self.snapshots
try:
self._follow()
except ControlError as e:
self.last_error = e.reason
if e.reason in _SLOW_RETRY_REASONS:
backoff = _RECONNECT_MAX_SECONDS
except Exception as e: # pylint: disable=broad-except
self.last_error = type(e).__name__
finally:
with self._lock:
self._connected = False
sock, self._sock = self._sock, None
if sock is not None:
try:
sock.close()
except OSError:
pass
if self.snapshots != snapshots:
# This connection got as far as the display's state: whatever
# ended it (a restart, most often), it was working, so the
# next try starts from the shortest wait again.
backoff = _RECONNECT_MIN_SECONDS
if self._stop.wait(backoff):
return
backoff = min(backoff * 2, _RECONNECT_MAX_SECONDS)
def _follow(self) -> None:
"""Subscribe, then read events until the connection ends. Raises ControlError."""
if not socket_supported():
raise ControlError('unsupported', 'no Unix sockets on this platform')
candidates = list(self._paths) if self._paths is not None else client_socket_paths()
if not candidates:
raise ControlError('disabled', 'the control socket is turned off')
request_id = str(uuid.uuid4())
payload = encode_message(Request(id=request_id, cmd=Command.STATE_SUBSCRIBE,
args={}).to_dict())
sock = _connect(candidates, time.monotonic() + self._connect_timeout)
self._sock = sock
try:
sock.settimeout(self._connect_timeout)
sock.sendall(payload)
# A read waits for the next event; the display sends one at least
# every keepalive, so this much silence means it is gone.
sock.settimeout(self._silence)
reader = FrameReader(MAX_MESSAGE_BYTES)
first = True
while not self._stop.is_set():
data = sock.recv(65536)
if not data:
raise ControlError('closed', 'the display closed the connection')
for line in reader.feed(data):
obj = decode_message(line)
if first:
response = Response.from_dict(obj)
if not response.ok:
error = response.error
raise ControlError(error.code if error else 'bad_response',
error.message if error else '')
self._store(dict(response.result or {}), full=True)
first = False
continue
event = StateEvent.from_dict(obj)
self._store(event.result, full=event.event == StateEventKind.STATE)
except socket.timeout:
raise ControlError('timeout', 'the display went quiet') from None
except ProtocolError as e:
raise ControlError('bad_response', e.message) from None
except OSError as e:
if self._stop.is_set():
return
raise ControlError('closed', str(e)) from None
def _store(self, result: Dict[str, Any], full: bool) -> None:
now = self._clock()
with self._lock:
if full and isinstance(result.get('state'), dict):
self._snapshot = result
self.snapshots += 1
elif (self._snapshot is not None
and result.get('epoch') == self._snapshot.get('epoch')):
# A tick: nothing changed but the render loop's liveness and
# the volatile keys (timestamps) the writers keep refreshing.
snap = dict(self._snapshot)
state = dict(snap.get('state') or {})
if result.get('version') == snap.get('version'):
_merge_volatile(state, result.get('volatile'))
loop = result.get('loop')
if isinstance(loop, dict):
state['loop'] = loop
snap['loop'] = loop
snap['state'] = state
snap['served_at'] = result.get('served_at', snap.get('served_at'))
self._snapshot = snap
else:
return # a tick before any state, or from another epoch
self._received = now
self._connected = True
+192 -3
View File
@@ -30,6 +30,11 @@ later the state stream). A few commands (:data:`AWAITED_COMMANDS`) are
answered only once the render thread has applied them, or with ``pending`` answered only once the render thread has applied them, or with ``pending``
when it has not within :data:`AWAIT_SECONDS`. when it has not within :data:`AWAIT_SECONDS`.
``state.subscribe`` is the one exception to "one response per request": its
response is followed, on the same connection, by :class:`StateEvent` lines
the display pushes until either side hangs up. Events carry ``event``
instead of ``ok``.
New commands are added within a protocol version: a display that does not New commands are added within a protocol version: a display that does not
know one answers ``unknown_command``, the client falls back, and ``hello`` know one answers ``unknown_command``, the client falls back, and ``hello``
lists the commands a display knows. The version changes only when the lists the commands a display knows. The version changes only when the
@@ -142,11 +147,14 @@ class Command:
ON_DEMAND_STATUS = 'on_demand.status' ON_DEMAND_STATUS = 'on_demand.status'
BRIGHTNESS_SET = 'brightness.set' BRIGHTNESS_SET = 'brightness.set'
PLUGIN_RELOAD = 'plugin.reload' PLUGIN_RELOAD = 'plugin.reload'
STATE_GET = 'state.get'
STATE_SUBSCRIBE = 'state.subscribe'
#: Every command version 1 defines, in the order ``hello`` reports them. #: Every command version 1 defines, in the order ``hello`` reports them.
#: ``brightness.set`` and ``plugin.reload`` came in stage 2, within version 1 #: ``brightness.set`` and ``plugin.reload`` came in stage 2, and ``state.get``
#: (see the module docstring on adding commands). #: and ``state.subscribe`` in stage 3, all within version 1 (see the module
#: docstring on adding commands).
COMMANDS: Tuple[str, ...] = ( COMMANDS: Tuple[str, ...] = (
Command.HELLO, Command.HELLO,
Command.PING, Command.PING,
@@ -155,6 +163,8 @@ COMMANDS: Tuple[str, ...] = (
Command.ON_DEMAND_STATUS, Command.ON_DEMAND_STATUS,
Command.BRIGHTNESS_SET, Command.BRIGHTNESS_SET,
Command.PLUGIN_RELOAD, Command.PLUGIN_RELOAD,
Command.STATE_GET,
Command.STATE_SUBSCRIBE,
) )
#: Commands that are queued for the render thread. #: Commands that are queued for the render thread.
@@ -173,6 +183,25 @@ AWAIT_SECONDS: Dict[str, float] = {
} }
AWAITED_COMMANDS = frozenset(AWAIT_SECONDS) AWAITED_COMMANDS = frozenset(AWAIT_SECONDS)
#: The state stream (stage 3). ``state.subscribe`` turns its connection into
#: a one-way stream of :class:`StateEvent` lines. Subscribers have their own
#: bound, separate from the short request connections, so they can never
#: take the slots a command needs.
MAX_SUBSCRIBERS = 4
#: A subscriber hears from the display at least this often: a ``state``
#: event when something changed, else a ``tick`` carrying the render loop's
#: liveness and the latest volatile timestamps. A client that has heard nothing for a few of these treats its
#: copy as unknown.
SUBSCRIBE_KEEPALIVE_SECONDS = 5.0
#: The shape of the ``state`` object in a state snapshot. Bumped only when a
#: field changes meaning; new fields are added within a schema.
STATE_SCHEMA = 1
#: The sections of a state snapshot, in the order they are documented.
STATE_SECTIONS: Tuple[str, ...] = ('display', 'on_demand', 'brightness', 'plugins', 'loop')
#: Brightness, in percent, as the display's hardware setting takes it. #: Brightness, in percent, as the display's hardware setting takes it.
MIN_BRIGHTNESS = 0 MIN_BRIGHTNESS = 0
MAX_BRIGHTNESS = 100 MAX_BRIGHTNESS = 100
@@ -477,8 +506,67 @@ class PluginReloadArgs:
return cls(plugin_id=plugin_id) return cls(plugin_id=plugin_id)
def _optional_version(args: Mapping[str, Any], key: str) -> Optional[int]:
value = args.get(key)
if value is None:
return None
if not _is_int(value) or value < 0:
raise ProtocolError(ErrorCode.INVALID_ARGS, f'{key} must be a non-negative integer')
return value
def _optional_epoch(args: Mapping[str, Any]) -> Optional[str]:
value = args.get('epoch')
if value is None or value == '':
return None
if not _valid_id(value):
raise ProtocolError(ErrorCode.INVALID_ARGS,
f'epoch must be a printable string of 1-{MAX_ID_LENGTH} characters')
return str(value)
@dataclass(frozen=True)
class StateGetArgs:
"""``state.get``: the display's state, as a versioned snapshot.
With ``since`` and the ``epoch`` it came from, the answer is only
``{changed: false, version, epoch, served_at, loop, volatile}`` while the
state is still at that version, so a poller that already has it is sent
no state -- only the latest values of the keys that do not count as a
change (``volatile``, see :class:`StateSnapshot`).
"""
since: Optional[int] = None
epoch: Optional[str] = None
def to_dict(self) -> Dict[str, Any]:
return {'since': self.since, 'epoch': self.epoch}
@classmethod
def from_dict(cls, args: Mapping[str, Any]) -> 'StateGetArgs':
return cls(since=_optional_version(args, 'since'), epoch=_optional_epoch(args))
@dataclass(frozen=True)
class StateSubscribeArgs:
"""``state.subscribe``: the snapshot now, then a push stream of changes.
The response is the snapshot ``state.get`` returns. After it the
connection carries only :class:`StateEvent` lines from the display: a
``state`` event whenever the state changes (always the latest version,
so a reader that falls behind skips versions instead of queueing them),
and a ``tick`` at least every :data:`SUBSCRIBE_KEEPALIVE_SECONDS`.
"""
def to_dict(self) -> Dict[str, Any]:
return {}
@classmethod
def from_dict(cls, args: Mapping[str, Any]) -> 'StateSubscribeArgs':
return cls()
CommandArgs = Union[HelloArgs, OnDemandStartArgs, OnDemandStopArgs, NoArgs, CommandArgs = Union[HelloArgs, OnDemandStartArgs, OnDemandStopArgs, NoArgs,
BrightnessSetArgs, PluginReloadArgs] BrightnessSetArgs, PluginReloadArgs, StateGetArgs, StateSubscribeArgs]
#: The arguments of a command that goes on the render thread's queue. #: The arguments of a command that goes on the render thread's queue.
QueuedArgs = Union[OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs] QueuedArgs = Union[OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs]
@@ -491,6 +579,8 @@ _ARG_TYPES: Dict[str, Any] = {
Command.ON_DEMAND_STATUS: NoArgs, Command.ON_DEMAND_STATUS: NoArgs,
Command.BRIGHTNESS_SET: BrightnessSetArgs, Command.BRIGHTNESS_SET: BrightnessSetArgs,
Command.PLUGIN_RELOAD: PluginReloadArgs, Command.PLUGIN_RELOAD: PluginReloadArgs,
Command.STATE_GET: StateGetArgs,
Command.STATE_SUBSCRIBE: StateSubscribeArgs,
} }
@@ -563,6 +653,105 @@ class PluginReloadResult(TypedDict):
modes: List[str] modes: List[str]
class LoopState(TypedDict):
"""``loop``: is the render loop still going round?
``heartbeat_age_seconds`` is the age of the render thread's last beat,
measured in memory by the display when it answered -- the same beat that
writes ``display-heartbeat.json``. None until the loop has drawn its first
frame. At ``stale_after`` or more the loop is stalled: the threshold
``/api/v3/health`` uses.
"""
heartbeat_age_seconds: Optional[float]
armed: bool
stale_after: float
class StateSnapshot(TypedDict, total=False):
"""The answer to ``state.get`` and ``state.subscribe``, and the
``result`` of a ``state`` event.
``version`` counts changes to the state within one ``epoch`` (one run of
the display process): a reader that sees a new epoch starts over.
``changed`` is False only for a ``state.get`` whose ``since`` is still
current, and then ``state`` is absent and ``volatile`` is there instead:
``{section: {key: value}}``, the current values of the keys the version
ignores (``display.last_updated``, ``on_demand.last_updated`` and
``remaining``, ``plugins.published_at``). A reader merges them into the
copy it has; they are how it can tell the writers are still publishing.
``served_at`` is the display's wall clock when it answered. ``loop`` is
measured at that moment, so it is also inside ``state``.
``state`` holds the sections in :data:`STATE_SECTIONS`:
* ``display``: what ``display_current_state`` holds (mode, plugin_id,
mode_index, total_modes, on_demand_active, is_display_active,
last_updated);
* ``on_demand``: what ``display_on_demand_state`` holds;
* ``brightness``: ``{brightness, panel_brightness, dimmed}``;
* ``plugins``: the plugin runtime snapshot (``plugin_runtime_snapshot``),
or None when there is none (or it was too large to send);
* ``loop``: :class:`LoopState`.
A section the display has not published yet is None.
"""
schema: int
version: int
epoch: str
pid: int
served_at: float
changed: bool
state: Dict[str, Any]
volatile: Dict[str, Dict[str, Any]]
loop: LoopState
class StateEventKind:
STATE = 'state' # result: a full StateSnapshot, the latest version
TICK = 'tick' # result: {version, epoch, pid, served_at, loop, volatile}; nothing changed
@dataclass(frozen=True)
class StateEvent:
"""One message the display pushes to a subscriber.
``{"v": 1, "id": "<the subscribe request's id>", "event": "state" | "tick",
"result": {...}}``. It has no ``ok``, which is how a reader tells it from
a response.
"""
id: str
event: str
result: Dict[str, Any]
v: int = PROTOCOL_VERSION
def to_dict(self) -> Dict[str, Any]:
return {'v': self.v, 'id': self.id, 'event': self.event, 'result': dict(self.result)}
@classmethod
def from_dict(cls, obj: Any) -> 'StateEvent':
"""Validate an event. Raises :class:`ProtocolError` (BAD_REQUEST)."""
if not isinstance(obj, dict):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'an event must be a JSON object')
version = obj.get('v')
if not _is_int(version):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'v must be an integer')
raw_id = obj.get('id')
if not isinstance(raw_id, str):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'id must be a string')
event = obj.get('event')
if event not in (StateEventKind.STATE, StateEventKind.TICK):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'event must be "state" or "tick"')
result = obj.get('result')
if not isinstance(result, dict):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'result must be a JSON object')
return cls(id=raw_id, event=event, result=result, v=version)
def is_event(obj: Any) -> bool:
"""Whether a decoded message is a pushed event rather than a response."""
return isinstance(obj, dict) and 'event' in obj and 'ok' not in obj
def negotiate_version(client_versions: Tuple[int, ...]) -> Optional[int]: def negotiate_version(client_versions: Tuple[int, ...]) -> Optional[int]:
"""The highest version both sides speak, or None.""" """The highest version both sides speak, or None."""
common = set(client_versions) & set(SUPPORTED_VERSIONS) common = set(client_versions) & set(SUPPORTED_VERSIONS)
+369 -16
View File
@@ -14,6 +14,13 @@ every kind of screen. An awaited command (``brightness.set``,
``plugin.reload``) carries a :class:`CommandOutcome` that the render thread ``plugin.reload``) carries a :class:`CommandOutcome` that the render thread
fills in; its connection thread waits for that, bounded, before answering. fills in; its connection thread waits for that, bounded, before answering.
The state stream (stage 3): the display publishes what it is doing into a
:class:`StateHub`, in memory, and ``state.get`` / ``state.subscribe`` read
it. A subscriber's connection gives back its request slot, takes one of
:data:`~src.ipc.contract.MAX_SUBSCRIBERS`, and is pushed the latest version
on every change plus a keepalive tick, from its own thread: publishing never
waits for a reader, and a reader that stops reading is dropped.
Robustness rules, because this runs inside the display process: Robustness rules, because this runs inside the display process:
* every connection has its own daemon thread, at most :data:`MAX_CLIENTS` at * every connection has its own daemon thread, at most :data:`MAX_CLIENTS` at
@@ -37,6 +44,7 @@ server checks them again: root, its own user, or a member of that group.
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
import os import os
import queue import queue
@@ -45,8 +53,9 @@ import stat
import struct import struct
import threading import threading
import time import time
import uuid
from dataclasses import dataclass, field from dataclasses import dataclass, field
from typing import Any, Callable, Dict, FrozenSet, List, Mapping, Optional from typing import Any, Callable, Dict, FrozenSet, Iterable, List, Mapping, Optional, Tuple
from src.ipc.contract import ( from src.ipc.contract import (
AWAIT_SECONDS, AWAIT_SECONDS,
@@ -55,8 +64,12 @@ from src.ipc.contract import (
DEFAULT_SOCKET_DIR, DEFAULT_SOCKET_DIR,
DEFAULT_SOCKET_PATH, DEFAULT_SOCKET_PATH,
MAX_MESSAGE_BYTES, MAX_MESSAGE_BYTES,
MAX_SUBSCRIBERS,
PROTOCOL_VERSION, PROTOCOL_VERSION,
QUEUED_COMMANDS, QUEUED_COMMANDS,
STATE_SCHEMA,
STATE_SECTIONS,
SUBSCRIBE_KEEPALIVE_SECONDS,
SUPPORTED_VERSIONS, SUPPORTED_VERSIONS,
AckResult, AckResult,
BrightnessSetArgs, BrightnessSetArgs,
@@ -72,6 +85,9 @@ from src.ipc.contract import (
QueuedArgs, QueuedArgs,
Request, Request,
Response, Response,
StateEvent,
StateEventKind,
StateGetArgs,
configured_socket_path, configured_socket_path,
decode_message, decode_message,
dev_socket_path, dev_socket_path,
@@ -181,6 +197,241 @@ class QueuedCommand:
self.outcome.fail(code, message) self.outcome.fail(code, message)
# -- the state stream (stage 3) ----------------------------------------------------------
#: How long a ``state.get`` keeps the display counting its readers as served
#: over the socket (:meth:`StateHub.readers_active`). A subscriber counts for
#: as long as it is connected.
READER_WINDOW_SECONDS = 60.0
#: Room kept for the envelope (``v``, ``id``, ``event``) around a snapshot,
#: within MAX_MESSAGE_BYTES.
_ENVELOPE_ROOM = 512
_MISSING = object()
LoopProbe = Callable[[], Mapping[str, Any]]
def _fingerprint(value: Optional[Mapping[str, Any]], volatile: Iterable[str]) -> Any:
"""What a section's version is judged on: the value minus its volatile keys
(timestamps that move on every publish without anything changing)."""
if value is None:
return None
skip = frozenset(volatile)
return {k: v for k, v in value.items() if k not in skip} if skip else dict(value)
def _volatile_values(sections: Mapping[str, Optional[Dict[str, Any]]],
volatile: Mapping[str, FrozenSet[str]]) -> Dict[str, Dict[str, Any]]:
"""``{section: {key: value}}``: the volatile keys each published section
has now. The sections are never mutated after publish (a publish swaps
in a new dict), so reading them outside the lock is safe."""
values: Dict[str, Dict[str, Any]] = {}
for name, keys in volatile.items():
value = sections.get(name)
if not keys or not isinstance(value, dict):
continue
present = {k: value[k] for k in keys if k in value}
if present:
values[name] = present
return values
def _unknown_loop() -> Dict[str, Any]:
return {'heartbeat_age_seconds': None, 'armed': False, 'stale_after': None}
def fit_snapshot(snapshot: Dict[str, Any]) -> Dict[str, Any]:
"""``snapshot``, or a copy without the plugin runtime section when the
message would be over MAX_MESSAGE_BYTES (hundreds of plugins). The
reader then falls back to the cache for that section only; ``truncated``
says which was left out."""
state = snapshot.get('state')
if not isinstance(state, dict) or state.get('plugins') is None:
return snapshot
try:
size = len(json.dumps(snapshot, separators=(',', ':'), ensure_ascii=True,
allow_nan=False))
except (TypeError, ValueError):
size = MAX_MESSAGE_BYTES
if size <= MAX_MESSAGE_BYTES - _ENVELOPE_ROOM:
return snapshot
logger.warning("State snapshot is %d bytes; sending it without the plugin runtime "
"section", size)
trimmed = dict(snapshot)
trimmed['state'] = dict(state, plugins=None)
trimmed['truncated'] = ['plugins']
return trimmed
class StateHub:
"""The display's live state, in memory, for ``state.get`` and ``state.subscribe``.
Writers publish whole sections (:meth:`publish`): the render thread
publishes ``display``, ``on_demand`` and ``brightness``, and the plugin
runtime publisher's thread publishes ``plugins``. Each section has one
writer. ``loop`` is not published: it is measured when a reader asks
(``loop_probe``), so it keeps ageing while the render thread is stuck.
The version goes up when a section's value changes, ignoring the keys
the publisher names as volatile (timestamps). Those keys still carry
news -- ``display.last_updated`` is the render thread's proof of life --
so the short ``changed: false`` answer, which is what a subscriber's
tick carries, has their current values in ``volatile``; a reader merges
them into its copy. Publishing never blocks on
a reader: the lock is held only to swap a dict reference and compare it,
and every socket write happens on the reader's own thread, outside it.
A reader that is slow gets the latest version when it next asks, not
every version in between.
"""
def __init__(self, loop_probe: Optional[LoopProbe] = None, *,
clock: Callable[[], float] = time.monotonic,
wall_clock: Callable[[], float] = time.time,
epoch: Optional[str] = None, pid: Optional[int] = None,
reader_window: float = READER_WINDOW_SECONDS):
self._cond = threading.Condition(threading.Lock())
self._sections: Dict[str, Optional[Dict[str, Any]]] = {}
self._fingerprints: Dict[str, Any] = {}
self._volatile: Dict[str, FrozenSet[str]] = {}
self._version = 0
self.epoch = epoch or uuid.uuid4().hex[:16]
self.pid = os.getpid() if pid is None else pid
self._loop_probe = loop_probe
self._clock = clock
self._wall_clock = wall_clock
self._reader_window = reader_window
self._last_read: Optional[float] = None
self._subscribers = 0
@property
def version(self) -> int:
return self._version
@property
def subscribers(self) -> int:
return self._subscribers
# -- writers -------------------------------------------------------------
def publish(self, section: str, value: Optional[Mapping[str, Any]],
volatile: Iterable[str] = ()) -> bool:
"""Store a section's latest value; True when that is a new version.
The value is copied (one level), so the caller may reuse its dict.
"""
stored = None if value is None else dict(value)
skip = frozenset(volatile)
fingerprint = _fingerprint(stored, skip)
with self._cond:
self._sections[section] = stored
self._volatile[section] = skip
if self._fingerprints.get(section, _MISSING) == fingerprint:
return False
self._fingerprints[section] = fingerprint
self._version += 1
self._cond.notify_all()
return True
def wake(self) -> None:
"""Wake every waiting reader (the server is closing)."""
with self._cond:
self._cond.notify_all()
# -- readers -------------------------------------------------------------
def loop(self) -> Dict[str, Any]:
"""The render loop's liveness now. Never raises."""
if self._loop_probe is None:
return _unknown_loop()
try:
return dict(self._loop_probe())
except Exception: # pylint: disable=broad-except
logger.debug("Render loop liveness probe failed", exc_info=True)
return _unknown_loop()
def snapshot(self, since: Optional[int] = None,
epoch: Optional[str] = None) -> Dict[str, Any]:
"""The :class:`~src.ipc.contract.StateSnapshot` now.
``since`` with this hub's ``epoch``, still the current version, gives
the short ``changed: false`` form, with ``volatile``: each section's
volatile keys at their latest values (the rest of the section is
what the reader already has).
"""
with self._cond:
version = self._version
sections = dict(self._sections)
volatile = dict(self._volatile)
loop = self.loop()
result: Dict[str, Any] = {
'schema': STATE_SCHEMA,
'version': version,
'epoch': self.epoch,
'pid': self.pid,
'served_at': self._wall_clock(),
'loop': loop,
}
if since is not None and epoch == self.epoch and since == version:
result['changed'] = False
result['volatile'] = _volatile_values(sections, volatile)
return result
state: Dict[str, Any] = {name: sections.get(name) for name in STATE_SECTIONS
if name != 'loop'}
state['loop'] = loop
result['changed'] = True
result['state'] = state
return result
def wait_for_change(self, version: int, timeout: float,
stop: Optional[threading.Event] = None) -> bool:
"""Block up to ``timeout`` for a version other than ``version``."""
with self._cond:
self._cond.wait_for(
lambda: self._version != version or (stop is not None and stop.is_set()),
timeout)
return self._version != version
# -- who is reading ------------------------------------------------------
def note_read(self) -> None:
self._last_read = self._clock()
def subscriber_joined(self) -> None:
with self._cond:
self._subscribers += 1
def subscriber_left(self) -> None:
with self._cond:
self._subscribers = max(0, self._subscribers - 1)
self._last_read = self._clock()
def readers_active(self) -> bool:
"""Is the socket serving state readers? A subscriber is connected, or a
``state.get`` came within the reader window. The display uses this to
write the cache copies of the same state less often."""
if self._subscribers > 0:
return True
last = self._last_read
return last is not None and self._clock() - last < self._reader_window
class _Slot:
"""A connection slot, released once (a subscriber gives its back early)."""
def __init__(self, semaphore: threading.BoundedSemaphore):
self._semaphore = semaphore
self._held = True
self._lock = threading.Lock()
def release(self) -> None:
with self._lock:
if self._held:
self._held = False
self._semaphore.release()
# -- peer credentials ------------------------------------------------------------------ # -- peer credentials ------------------------------------------------------------------
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -315,8 +566,14 @@ class ControlServer:
message_timeout: float = MESSAGE_TIMEOUT_SECONDS, message_timeout: float = MESSAGE_TIMEOUT_SECONDS,
idle_timeout: float = IDLE_TIMEOUT_SECONDS, idle_timeout: float = IDLE_TIMEOUT_SECONDS,
check_peer: bool = True, check_peer: bool = True,
await_seconds: Optional[Mapping[str, float]] = None): await_seconds: Optional[Mapping[str, float]] = None,
state_hub: Optional[StateHub] = None,
max_subscribers: int = MAX_SUBSCRIBERS,
keepalive: float = SUBSCRIBE_KEEPALIVE_SECONDS):
self.path = path self.path = path
self.state_hub = state_hub
self._subscriber_slots = threading.BoundedSemaphore(max_subscribers)
self._keepalive = keepalive
self._await_seconds: Dict[str, float] = dict(AWAIT_SECONDS) self._await_seconds: Dict[str, float] = dict(AWAIT_SECONDS)
if await_seconds: if await_seconds:
self._await_seconds.update(await_seconds) self._await_seconds.update(await_seconds)
@@ -378,6 +635,8 @@ class ControlServer:
"""Stop accepting and remove the socket file (only if it is still ours).""" """Stop accepting and remove the socket file (only if it is still ours)."""
self._stopping.set() self._stopping.set()
self._close_socket() self._close_socket()
if self.state_hub is not None:
self.state_hub.wake() # subscribers see _stopping and hang up
thread = self._thread thread = self._thread
if thread is not None and thread is not threading.current_thread(): if thread is not None and thread is not threading.current_thread():
thread.join(timeout=2.0) thread.join(timeout=2.0)
@@ -549,6 +808,7 @@ class ControlServer:
def _serve(self, conn: socket.socket) -> None: def _serve(self, conn: socket.socket) -> None:
"""One connection: authenticate, then answer requests until it ends.""" """One connection: authenticate, then answer requests until it ends."""
slot = _Slot(self._slots)
try: try:
conn.settimeout(self._io_timeout) conn.settimeout(self._io_timeout)
peer = peer_credentials(conn) peer = peer_credentials(conn)
@@ -557,7 +817,7 @@ class ControlServer:
"this user or group %s", peer.pid, peer.uid, peer.gid, self._group) "this user or group %s", peer.pid, peer.uid, peer.gid, self._group)
self._send(conn, Response.failure(None, ErrorCode.FORBIDDEN, 'not permitted')) self._send(conn, Response.failure(None, ErrorCode.FORBIDDEN, 'not permitted'))
return return
self._read_requests(conn, peer) self._read_requests(conn, peer, slot)
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Control socket connection failed") logger.exception("Control socket connection failed")
finally: finally:
@@ -565,7 +825,7 @@ class ControlServer:
conn.close() conn.close()
except OSError: except OSError:
pass pass
self._slots.release() slot.release()
def _peer_ok(self, peer: PeerCredentials) -> bool: def _peer_ok(self, peer: PeerCredentials) -> bool:
groups = None groups = None
@@ -573,7 +833,8 @@ class ControlServer:
groups = process_groups(peer.pid) groups = process_groups(peer.pid)
return peer_allowed(peer, self._own_uid, self._group, groups) return peer_allowed(peer, self._own_uid, self._group, groups)
def _read_requests(self, conn: socket.socket, peer: Optional[PeerCredentials]) -> None: def _read_requests(self, conn: socket.socket, peer: Optional[PeerCredentials],
slot: Optional[_Slot] = None) -> None:
reader = FrameReader(MAX_MESSAGE_BYTES) reader = FrameReader(MAX_MESSAGE_BYTES)
idle_since = time.monotonic() idle_since = time.monotonic()
message_started: Optional[float] = None message_started: Optional[float] = None
@@ -598,7 +859,13 @@ class ControlServer:
self._send(conn, Response.failure(None, e.code, e.message)) self._send(conn, Response.failure(None, e.code, e.message))
return # can't find the next message boundary: hang up return # can't find the next message boundary: hang up
for line in lines: for line in lines:
if not self._send(conn, self.handle_line(line, peer)): response, cmd = self._handle(line, peer)
if cmd == Command.STATE_SUBSCRIBE and response.ok:
# The connection becomes a one-way stream; anything the
# client sent after the subscribe is ignored.
self._subscribe(conn, response, slot)
return
if not self._send(conn, response):
return return
if reader.pending: if reader.pending:
if message_started is None or lines: if message_started is None or lines:
@@ -624,20 +891,91 @@ class ControlServer:
# -- requests -------------------------------------------------------------------- # -- requests --------------------------------------------------------------------
def handle_line(self, line: bytes, peer: Optional[PeerCredentials] = None) -> Response: def handle_line(self, line: bytes, peer: Optional[PeerCredentials] = None) -> Response:
"""Answer one request line. Never raises.""" """Answer one request line. Never raises.
A ``state.subscribe`` answered here gets its snapshot only; the
stream that follows needs a connection (``_read_requests``).
"""
return self._handle(line, peer)[0]
def _handle(self, line: bytes,
peer: Optional[PeerCredentials]) -> Tuple[Response, Optional[str]]:
"""The response to one line, and the command it answered (when known)."""
request_id: Optional[str] = None request_id: Optional[str] = None
cmd: Optional[str] = None
try: try:
obj = decode_message(line) obj = decode_message(line)
raw_id = obj.get('id') raw_id = obj.get('id')
request_id = raw_id if isinstance(raw_id, str) and len(raw_id) <= 128 else None request_id = raw_id if isinstance(raw_id, str) and len(raw_id) <= 128 else None
request = Request.from_dict(obj) request = Request.from_dict(obj)
request_id = request.id request_id = request.id
return self._dispatch(request, peer) cmd = request.cmd
return self._dispatch(request, peer), cmd
except ProtocolError as e: except ProtocolError as e:
return Response.failure(e.request_id or request_id, e.code, e.message) return Response.failure(e.request_id or request_id, e.code, e.message), cmd
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Control socket handler failed") logger.exception("Control socket handler failed")
return Response.failure(request_id, ErrorCode.INTERNAL, 'internal error') return Response.failure(request_id, ErrorCode.INTERNAL, 'internal error'), cmd
# -- the state stream ----------------------------------------------------------
def _subscribe(self, conn: socket.socket, response: Response,
slot: Optional[_Slot]) -> None:
"""Answer a ``state.subscribe`` and push state events until it ends.
Subscribers have their own bound (MAX_SUBSCRIBERS) and give their
request slot back, so a few browsers watching never use up the slots
commands need. Everything here runs on this connection's thread: a
reader that does not keep up only stalls its own sends, and one that
stops reading for a whole IO timeout is dropped. The render thread
only ever publishes into the hub.
"""
hub = self.state_hub
if hub is None or not self._subscriber_slots.acquire(blocking=False):
self._send(conn, Response.failure(response.id, ErrorCode.BUSY,
'too many state subscribers', v=response.v))
return
if slot is not None:
slot.release()
hub.subscriber_joined()
try:
if not self._send(conn, response):
return
result = response.result or {}
version = result.get('version', -1)
sub_id = response.id or ''
logger.debug("Control socket: state subscriber joined at version %s", version)
while not self._stopping.is_set():
hub.wait_for_change(version, self._keepalive, self._stopping)
if self._stopping.is_set():
return
snap = hub.snapshot(since=version, epoch=hub.epoch)
if snap.get('changed'):
version = snap['version']
event = StateEvent(sub_id, StateEventKind.STATE, fit_snapshot(snap),
v=response.v)
else:
event = StateEvent(sub_id, StateEventKind.TICK, snap, v=response.v)
if not self._send_event(conn, event):
return
finally:
hub.subscriber_left()
self._subscriber_slots.release()
def _send_event(self, conn: socket.socket, event: StateEvent) -> bool:
try:
data = encode_message(event.to_dict())
except ProtocolError as e:
logger.error("Control socket state event not sent: %s", e.message)
return False
try:
conn.sendall(data)
return True
except socket.timeout:
logger.info("Control socket: dropping a state subscriber that stopped reading")
return False
except OSError:
return False
def _dispatch(self, request: Request, peer: Optional[PeerCredentials]) -> Response: def _dispatch(self, request: Request, peer: Optional[PeerCredentials]) -> Response:
if request.cmd == Command.HELLO: if request.cmd == Command.HELLO:
@@ -678,6 +1016,18 @@ class ControlServer:
v=request.v) v=request.v)
return Response.success(request.id, self._status_provider(), v=request.v) return Response.success(request.id, self._status_provider(), v=request.v)
if request.cmd in (Command.STATE_GET, Command.STATE_SUBSCRIBE):
hub = self.state_hub
if hub is None:
return Response.failure(request.id, ErrorCode.INTERNAL, 'no state available',
v=request.v)
if isinstance(args, StateGetArgs):
hub.note_read()
snap = hub.snapshot(since=args.since, epoch=args.epoch)
else:
snap = hub.snapshot()
return Response.success(request.id, fit_snapshot(snap), v=request.v)
if request.cmd in QUEUED_COMMANDS and isinstance(args, ( if request.cmd in QUEUED_COMMANDS and isinstance(args, (
OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs)): OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs)):
awaited = request.cmd in AWAITED_COMMANDS awaited = request.cmd in AWAITED_COMMANDS
@@ -728,22 +1078,25 @@ class ControlServer:
def start_control_server(status_provider: Optional[StatusProvider] = None, def start_control_server(status_provider: Optional[StatusProvider] = None,
cache_dir: Optional[str] = None, cache_dir: Optional[str] = None,
environ: Optional[Mapping[str, str]] = None) -> Optional[ControlServer]: environ: Optional[Mapping[str, str]] = None,
state_hub: Optional[StateHub] = None) -> Optional[ControlServer]:
"""Start the display's control socket, or return None when it can't run. """Start the display's control socket, or return None when it can't run.
None covers Windows, ``LEDMATRIX_CONTROL_SOCKET=off`` and any failure to None covers Windows, ``LEDMATRIX_CONTROL_SOCKET=off`` and any failure to
bind; in every case the web interface falls back to the file mailbox. bind; in every case the web interface falls back to the file mailbox
and to the cache keys the display still writes.
""" """
path = server_socket_path(environ) path = server_socket_path(environ)
if path is None: if path is None:
logger.debug("Control socket disabled or unsupported here; using the file mailbox only") logger.debug("Control socket disabled or unsupported here; using the file mailbox only")
return None return None
server = ControlServer(path, status_provider, resolve_socket_group(cache_dir)) server = ControlServer(path, status_provider, resolve_socket_group(cache_dir),
state_hub=state_hub)
return server if server.start() else None return server if server.start() else None
__all__ = [ __all__ = [
'CommandOutcome', 'ControlServer', 'PeerCredentials', 'QueuedCommand', 'StatusProvider', 'CommandOutcome', 'ControlServer', 'PeerCredentials', 'QueuedCommand', 'StateHub',
'peer_allowed', 'peer_credentials', 'process_groups', 'resolve_socket_group', 'StatusProvider', 'fit_snapshot', 'peer_allowed', 'peer_credentials', 'process_groups',
'server_socket_path', 'start_control_server', 'PROTOCOL_VERSION', 'resolve_socket_group', 'server_socket_path', 'start_control_server', 'PROTOCOL_VERSION',
] ]
+3 -2
View File
@@ -21,6 +21,7 @@ from PIL.PngImagePlugin import PngInfo
from requests.adapters import HTTPAdapter from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry from urllib3.util.retry import Retry
from src.common.api_helper import DEFAULT_HTTP_HEADERS from src.common.api_helper import DEFAULT_HTTP_HEADERS
from src.common.json_body import response_json
from src.common.logo_helper import MAX_LOGO_BYTES from src.common.logo_helper import MAX_LOGO_BYTES
from src.common.permission_utils import ( from src.common.permission_utils import (
ensure_directory_permissions, ensure_directory_permissions,
@@ -481,7 +482,7 @@ class LogoDownloader:
logger.info(f"Fetching team data for {league} from ESPN API...") logger.info(f"Fetching team data for {league} from ESPN API...")
response = self.session.get(api_url, params={'limit':1000},headers=self.headers, timeout=self.request_timeout) response = self.session.get(api_url, params={'limit':1000},headers=self.headers, timeout=self.request_timeout)
response.raise_for_status() response.raise_for_status()
data: Dict = response.json() data: Dict = response_json(response)
logger.info(f"Successfully fetched team data for {league}") logger.info(f"Successfully fetched team data for {league}")
return data return data
@@ -505,7 +506,7 @@ class LogoDownloader:
logger.info(f"Fetching team data for team {team_id} in {league} from ESPN API...") logger.info(f"Fetching team data for team {team_id} in {league} from ESPN API...")
response = self.session.get(f"{api_url}/{team_id}", headers=self.headers, timeout=self.request_timeout) response = self.session.get(f"{api_url}/{team_id}", headers=self.headers, timeout=self.request_timeout)
response.raise_for_status() response.raise_for_status()
data: Dict = response.json() data: Dict = response_json(response)
logger.info(f"Successfully fetched team data for {team_id} in {league}") logger.info(f"Successfully fetched team data for {team_id} in {league}")
return data return data
+33 -2
View File
@@ -3,15 +3,46 @@ LEDMatrix Plugin System
This module provides the core plugin infrastructure for the LEDMatrix project. This module provides the core plugin infrastructure for the LEDMatrix project.
It enables dynamic loading, management, and discovery of display plugins. It enables dynamic loading, management, and discovery of display plugins.
BasePlugin and PluginManager are imported on first use (PEP 562), not when
the package is imported: the web interface imports several submodules
(store_manager, schema_manager, ...) and never needs PluginManager, which
pulls in the loader, executor and the shared helpers behind them.
``from src.plugin_system import BasePlugin`` works as before and returns the
same class.
""" """
import importlib
from typing import TYPE_CHECKING, Any, Dict, List, Tuple
__version__ = "1.0.0" __version__ = "1.0.0"
from .base_plugin import BasePlugin if TYPE_CHECKING:
from .plugin_manager import PluginManager from .base_plugin import BasePlugin
from .plugin_manager import PluginManager
#: Exported name -> (module it lives in, attribute name there).
_LAZY: Dict[str, Tuple[str, str]] = {
'BasePlugin': ('src.plugin_system.base_plugin', 'BasePlugin'),
'PluginManager': ('src.plugin_system.plugin_manager', 'PluginManager'),
}
__all__ = [ __all__ = [
'BasePlugin', 'BasePlugin',
'PluginManager', 'PluginManager',
] ]
def __getattr__(name: str) -> Any:
"""Import an exported name on first access (PEP 562); see src.common."""
try:
module_name, attr = _LAZY[name]
except KeyError:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None
value = getattr(importlib.import_module(module_name), attr) # nosemgrep: python.lang.security.audit.non-literal-import.non-literal-import -- module_name comes from the fixed _LAZY table
globals()[name] = value
return value
def __dir__() -> List[str]:
return sorted(set(globals()) | set(__all__))
+4 -1
View File
@@ -92,7 +92,10 @@ class PluginExecutor:
with plugin_scope(plugin_id): with plugin_scope(plugin_id):
result_container['value'] = operation() result_container['value'] = operation()
result_container['completed'] = True result_container['completed'] = True
except Exception as e: except BaseException as e: # pylint: disable=broad-except
# asyncio.CancelledError and SystemExit too: uncaught, one
# ended this thread with 'completed' unset, and an operation
# that failed at once was reported as timing out.
result_container['exception'] = e result_container['exception'] = e
result_container['completed'] = True result_container['completed'] = True
+75 -4
View File
@@ -199,9 +199,22 @@ def contained_plugin_dir(plugin_dir: Path, plugins_dir: Path) -> Optional[str]:
name that came out of ``os.scandir()`` on the trusted root carries no name that came out of ``os.scandir()`` on the trusted root carries no
taint, which is a real containment guarantee (and one CodeQL's taint, which is a real containment guarantee (and one CodeQL's
path-injection query can follow), not a string sanitiser. path-injection query can follow), not a string sanitiser.
The entry looked for is the one ``plugin_dir`` itself names when it sits
directly in ``plugins_dir``: for a dev plugin symlinked in under its id,
the link's name. Resolving the link first and looking for the target's
folder name refused ``plugins/foo -> ~/.ledmatrix-dev-plugins/ledmatrix-foo``
(what ``dev_plugin_setup.sh link-github foo <url>`` makes), so the plugin
never loaded. Any other path is resolved and matched by its final name,
as before.
""" """
plugin_dir_real = os.path.realpath(str(plugin_dir))
plugins_dir_real = os.path.realpath(str(plugins_dir)) plugins_dir_real = os.path.realpath(str(plugins_dir))
plugin_dir_abs = os.path.abspath(str(plugin_dir))
if os.path.realpath(os.path.dirname(plugin_dir_abs)) == plugins_dir_real:
matched_name = find_trusted_subdir(plugins_dir_real, os.path.basename(plugin_dir_abs))
if matched_name is not None:
return os.path.join(plugins_dir_real, matched_name)
plugin_dir_real = os.path.realpath(str(plugin_dir))
matched_name = find_trusted_subdir(plugins_dir_real, os.path.basename(plugin_dir_real)) matched_name = find_trusted_subdir(plugins_dir_real, os.path.basename(plugin_dir_real))
if matched_name is None: if matched_name is None:
return None return None
@@ -243,6 +256,10 @@ class PluginLoader:
self.logger = logger or get_logger(__name__) self.logger = logger or get_logger(__name__)
self._loaded_modules: Dict[str, Any] = {} self._loaded_modules: Dict[str, Any] = {}
self._plugin_module_registry: Dict[str, set] = {} # Maps plugin_id to set of module names self._plugin_module_registry: Dict[str, set] = {} # Maps plugin_id to set of module names
# plugin_id -> {dotted name: module} for the modules of the plugin's
# own packages (``providers.feed``). They keep their names while the
# plugin runs and are dropped with it; see _iter_plugin_submodules.
self._plugin_submodules: Dict[str, Dict[str, Any]] = {}
# Lock to serialize module loading when plugins share module names # Lock to serialize module loading when plugins share module names
# (e.g., scroll_display.py, game_renderer.py across sport plugins). # (e.g., scroll_display.py, game_renderer.py across sport plugins).
# During exec_module, bare-name sub-modules temporarily appear in # During exec_module, bare-name sub-modules temporarily appear in
@@ -449,6 +466,45 @@ class PluginLoader:
continue continue
return result return result
@staticmethod
def _iter_plugin_submodules(
plugin_dir: Path, before_keys: set
) -> list:
"""Return dotted-name modules from plugin_dir added after before_keys.
The modules of a package the plugin ships (``providers.feed`` from
``providers/feed.py``). _iter_plugin_bare_modules skips them, so the
bare ``providers`` was namespaced and dropped on unload while
``providers.feed`` stayed in sys.modules: a reload after a store update
imported a fresh ``providers`` and then got the old ``feed`` back from
the cache, running the new manager.py against the old helpers until the
display restarted.
A module counts when its ``__file__`` -- or, for a namespace package,
which has none, every ``__path__`` entry -- is inside plugin_dir, so a
library the plugin imports (``requests.adapters``) never does.
Returns a list of (mod_name, module) tuples.
"""
resolved_dir = plugin_dir.resolve()
result = []
for key in set(sys.modules.keys()) - before_keys:
if "." not in key:
continue
mod = sys.modules.get(key)
if mod is None:
continue
mod_file = getattr(mod, "__file__", None)
locations = [mod_file] if mod_file else list(getattr(mod, "__path__", None) or [])
if not locations:
continue
try:
if all(Path(loc).resolve().is_relative_to(resolved_dir) for loc in locations):
result.append((key, mod))
except (ValueError, TypeError, OSError):
continue
return result
def _evict_stale_bare_modules(self, plugin_dir: Path) -> dict: def _evict_stale_bare_modules(self, plugin_dir: Path) -> dict:
"""Temporarily remove bare-name sys.modules entries from other plugins. """Temporarily remove bare-name sys.modules entries from other plugins.
@@ -527,6 +583,13 @@ class PluginLoader:
# Track for cleanup during unload # Track for cleanup during unload
self._plugin_module_registry[plugin_id] = namespaced_names self._plugin_module_registry[plugin_id] = namespaced_names
# The modules of the plugin's own packages keep their dotted names
# while it runs -- as they always have, so the package and its
# children stay a matching set in sys.modules -- and are dropped
# with the plugin by unregister_plugin_modules().
self._plugin_submodules[plugin_id] = dict(
self._iter_plugin_submodules(plugin_dir, before_keys))
if namespaced_names: if namespaced_names:
self.logger.info( self.logger.info(
"Namespace-isolated %d module(s) for plugin %s", "Namespace-isolated %d module(s) for plugin %s",
@@ -537,10 +600,16 @@ class PluginLoader:
"""Remove namespaced sub-modules and cached module for a plugin from sys.modules. """Remove namespaced sub-modules and cached module for a plugin from sys.modules.
Called by PluginManager during unload to clean up all module entries Called by PluginManager during unload to clean up all module entries
that were created when the plugin was loaded. that were created when the plugin was loaded, including the dotted
modules of its packages. A dotted name is dropped only while it still
holds this plugin's module: the name is not namespaced, so another
plugin may have put its own there since.
""" """
for ns_name in self._plugin_module_registry.pop(plugin_id, set()): for ns_name in self._plugin_module_registry.pop(plugin_id, set()):
sys.modules.pop(ns_name, None) sys.modules.pop(ns_name, None)
for name, mod in self._plugin_submodules.pop(plugin_id, {}).items():
if sys.modules.get(name) is mod:
sys.modules.pop(name, None)
self._loaded_modules.pop(plugin_id, None) self._loaded_modules.pop(plugin_id, None)
def load_module( def load_module(
@@ -646,11 +715,13 @@ class PluginLoader:
if evicted_name not in sys.modules: if evicted_name not in sys.modules:
sys.modules[evicted_name] = evicted_mod sys.modules[evicted_name] = evicted_mod
# Clean up the partially-initialized main module and any # Clean up the partially-initialized main module and any
# bare-name sub-modules that were added during exec_module # bare-name or package sub-modules that were added during
# so they don't leak into subsequent plugin loads. # exec_module so they don't leak into subsequent plugin loads.
sys.modules.pop(module_name, None) sys.modules.pop(module_name, None)
for key, _ in self._iter_plugin_bare_modules(plugin_dir, before_keys): for key, _ in self._iter_plugin_bare_modules(plugin_dir, before_keys):
sys.modules.pop(key, None) sys.modules.pop(key, None)
for key, _ in self._iter_plugin_submodules(plugin_dir, before_keys):
sys.modules.pop(key, None)
raise raise
self._loaded_modules[plugin_id] = module self._loaded_modules[plugin_id] = module
+10 -4
View File
@@ -1163,7 +1163,7 @@ class PluginManager:
def _record_update_failure( def _record_update_failure(
self, self,
plugin_id: str, plugin_id: str,
exc: Optional[Exception] = None, exc: Optional[BaseException] = None,
log: bool = True, log: bool = True,
count_failure: bool = True, count_failure: bool = True,
) -> None: ) -> None:
@@ -1187,7 +1187,7 @@ class PluginManager:
""" """
failure_time = time.time() failure_time = time.time()
if exc is not None: if exc is not None:
err: Exception = exc err: BaseException = exc
error_type = type(exc).__name__ error_type = type(exc).__name__
else: else:
err = Exception(f"Plugin {plugin_id} execution failed (timeout or executor error)") err = Exception(f"Plugin {plugin_id} execution failed (timeout or executor error)")
@@ -1653,7 +1653,7 @@ class PluginManager:
finish_guard = threading.Lock() finish_guard = threading.Lock()
finished = {'done': False} finished = {'done': False}
def _finish(success: bool, exc: Optional[Exception] = None) -> None: def _finish(success: bool, exc: Optional[BaseException] = None) -> None:
with finish_guard: with finish_guard:
if finished['done']: if finished['done']:
return return
@@ -1727,7 +1727,13 @@ class PluginManager:
self.resource_monitor.monitor_call(plugin_id, plugin_instance.update) self.resource_monitor.monitor_call(plugin_id, plugin_instance.update)
else: else:
plugin_instance.update() plugin_instance.update()
except Exception as exc: except BaseException as exc: # pylint: disable=broad-except
# BaseException, not just Exception: asyncio.CancelledError
# and SystemExit derive from it. Either one skipped _finish,
# so the plugin kept its lock and stayed RUNNING for good --
# never rescheduled, and every display() skipped as busy.
# Re-raised for the executor, which reports it as this
# update's failure.
_finish(False, exc=exc) _finish(False, exc=exc)
raise raise
else: else:
+117 -9
View File
@@ -36,6 +36,13 @@ tmpfs. A missing heartbeat (dev server, emulator, Windows, a display still
starting up) or one from another process (a display restarted after a starting up) or one from another process (a display restarted after a
watchdog kill) says nothing, and the snapshot is judged on its own. watchdog kill) says nothing, and the snapshot is judged on its own.
The control socket. Where the display serves its state stream (stage 3,
docs/IPC_CONTROL_SOCKET.md), every tick also hands the snapshot to it, in
memory, and the web interface reads it there first
(``view_from_socket_state``, judged by the same rules). While the socket
serves those readers, the cache copy is their fallback and an unchanged
snapshot is rewritten every ``RELAXED_REFRESH_INTERVAL`` instead.
A dead publisher. systemd removes the heartbeat's directory when the A dead publisher. systemd removes the heartbeat's directory when the
service stops, so after a watchdog kill there is no heartbeat to go stale. service stops, so after a watchdog kill there is no heartbeat to go stale.
The reader then asks whether the snapshot's ``pid`` still exists (POSIX The reader then asks whether the snapshot's ``pid`` still exists (POSIX
@@ -48,7 +55,7 @@ import math
import os import os
import threading import threading
import time import time
from dataclasses import dataclass, field from dataclasses import dataclass, field, replace
from typing import Any, Callable, Dict, Optional from typing import Any, Callable, Dict, Optional
from src import display_watchdog from src import display_watchdog
@@ -75,6 +82,15 @@ TICK_INTERVAL = 5.0
#: A snapshot older than this is stale: three missed refreshes. #: A snapshot older than this is stale: three missed refreshes.
STALE_AFTER = 3 * REFRESH_INTERVAL STALE_AFTER = 3 * REFRESH_INTERVAL
#: The refresh while the control socket serves the web interface's readers
#: (``StateHub.readers_active``). The cache copy is then only their fallback,
#: so an unchanged snapshot is rewritten half as often; the snapshot says so
#: in its own ``refresh_interval`` and ``stale_after``.
RELAXED_REFRESH_INTERVAL = 2 * REFRESH_INTERVAL
#: The control socket's section for this snapshot (``state.plugins``).
STATE_SECTION = "plugins"
#: Bounds on a published ``stale_after``, so a corrupt value can make a #: Bounds on a published ``stale_after``, so a corrupt value can make a
#: reader neither trust a dead display for hours nor distrust a live one. #: reader neither trust a dead display for hours nor distrust a live one.
_STALE_AFTER_MIN = 30.0 _STALE_AFTER_MIN = 30.0
@@ -140,7 +156,8 @@ def summarize_error(error_info: Optional[Dict[str, Any]]) -> Optional[Dict[str,
def build_runtime_snapshot(state_manager: Any, *, started_at: float, def build_runtime_snapshot(state_manager: Any, *, started_at: float,
now: Optional[float] = None, now: Optional[float] = None,
running: bool = True) -> Dict[str, Any]: running: bool = True,
refresh_interval: float = REFRESH_INTERVAL) -> Dict[str, Any]:
"""The snapshot for ``state_manager`` (a plugin_state.PluginStateManager). """The snapshot for ``state_manager`` (a plugin_state.PluginStateManager).
A stopped snapshot (``running=False``) lists no plugins: nothing is A stopped snapshot (``running=False``) lists no plugins: nothing is
@@ -162,8 +179,8 @@ def build_runtime_snapshot(state_manager: Any, *, started_at: float,
"running": running, "running": running,
"published_at": time.time() if now is None else now, "published_at": time.time() if now is None else now,
"started_at": started_at, "started_at": started_at,
"refresh_interval": REFRESH_INTERVAL, "refresh_interval": refresh_interval,
"stale_after": STALE_AFTER, "stale_after": 3 * refresh_interval,
"pid": os.getpid(), "pid": os.getpid(),
"plugins": plugins, "plugins": plugins,
} }
@@ -197,30 +214,87 @@ class PluginRuntimePublisher:
self._tick_lock = threading.Lock() self._tick_lock = threading.Lock()
self._stop = threading.Event() self._stop = threading.Event()
self._thread: Optional[threading.Thread] = None self._thread: Optional[threading.Thread] = None
# The control socket's state stream (src/ipc/server.StateHub), when
# the display serves one: every tick also hands it the snapshot, in
# memory, and the cache refresh relaxes while it has readers.
self._hub: Any = None
self._hub_change: Optional[int] = None
self._hub_snapshot: Optional[Dict[str, Any]] = None
self.relaxed_refresh_interval = RELAXED_REFRESH_INTERVAL
def _write(self, running: bool) -> None: def attach_hub(self, hub: Any) -> None:
"""Also publish to the control socket's state hub, starting now."""
with self._tick_lock:
self._hub = hub
self._hub_change = None
self._hub_snapshot = None
try:
self._push_to_hub(self.state_manager.change_count)
except Exception as err: # never let reporting break the display
logger.debug("Could not publish the plugin runtime state: %s", err,
exc_info=True)
def _push_to_hub(self, change: int) -> None:
"""The snapshot to the state hub: rebuilt when the state machine
changed, otherwise the last one with a new ``published_at``, which
the hub does not count as a new version. In memory, every tick, so
the socket's copy is never more than a tick old."""
hub = self._hub
if hub is None:
return
now = self._wall_clock()
if self._hub_snapshot is None or change != self._hub_change:
snapshot = build_runtime_snapshot(self.state_manager, started_at=self.started_at, snapshot = build_runtime_snapshot(self.state_manager, started_at=self.started_at,
now=self._wall_clock(), running=running) now=now)
else:
snapshot = dict(self._hub_snapshot, published_at=now)
hub.publish(STATE_SECTION, snapshot, volatile=("published_at",))
self._hub_snapshot = snapshot
self._hub_change = change
def _cache_refresh_interval(self) -> float:
"""The cache refresh: relaxed while the socket serves the readers."""
hub = self._hub
try:
if hub is not None and hub.readers_active():
return self.relaxed_refresh_interval
except Exception: # pylint: disable=broad-except
# The normal interval is the safe answer: it only writes more.
logger.debug("State hub readers_active() failed; using the normal refresh", exc_info=True)
return self.refresh_interval
def _write(self, running: bool, refresh_interval: Optional[float] = None) -> None:
snapshot = build_runtime_snapshot(
self.state_manager, started_at=self.started_at, now=self._wall_clock(),
running=running,
refresh_interval=self.refresh_interval if refresh_interval is None
else refresh_interval)
self.cache_manager.set(PLUGIN_RUNTIME_KEY, snapshot) self.cache_manager.set(PLUGIN_RUNTIME_KEY, snapshot)
def tick(self) -> bool: def tick(self) -> bool:
"""Publish if something changed (throttled) or the refresh is due. """Publish if something changed (throttled) or the refresh is due.
True if a snapshot was written.""" True if a snapshot was written to the cache."""
with self._tick_lock: with self._tick_lock:
try: try:
change = self.state_manager.change_count change = self.state_manager.change_count
try:
self._push_to_hub(change)
except Exception as err: # the cache copy still goes out below
logger.debug("Could not publish the plugin runtime state: %s", err,
exc_info=True)
now = self._clock() now = self._clock()
refresh = self._cache_refresh_interval()
since = None if self._last_attempt is None else now - self._last_attempt since = None if self._last_attempt is None else now - self._last_attempt
if since is not None: if since is not None:
if change == self._published_change: if change == self._published_change:
if since < self.refresh_interval: if since < refresh:
return False return False
elif since < self.min_interval: elif since < self.min_interval:
return False return False
# Stamp the attempt before writing: a cache that keeps failing # Stamp the attempt before writing: a cache that keeps failing
# is retried at the throttled rate, not on every tick. # is retried at the throttled rate, not on every tick.
self._last_attempt = now self._last_attempt = now
self._write(running=True) self._write(running=True, refresh_interval=refresh)
self._published_change = change self._published_change = change
return True return True
except Exception as err: # never let reporting break the display except Exception as err: # never let reporting break the display
@@ -317,6 +391,9 @@ class PluginRuntimeView:
plugins: Dict[str, Dict[str, Any]] = field(default_factory=dict) plugins: Dict[str, Dict[str, Any]] = field(default_factory=dict)
#: Age of the render loop's heartbeat, when it was taken into account. #: Age of the render loop's heartbeat, when it was taken into account.
heartbeat_age_seconds: Optional[float] = None heartbeat_age_seconds: Optional[float] = None
#: Where the snapshot came from: ``cache`` (the shared cache file and the
#: heartbeat file) or ``socket`` (the control socket's state stream).
source: str = "cache"
@property @property
def live(self) -> bool: def live(self) -> bool:
@@ -348,6 +425,7 @@ class PluginRuntimeView:
"stale_after": self.stale_after, "stale_after": self.stale_after,
"heartbeat_age_seconds": (None if self.heartbeat_age_seconds is None "heartbeat_age_seconds": (None if self.heartbeat_age_seconds is None
else round(self.heartbeat_age_seconds, 1)), else round(self.heartbeat_age_seconds, 1)),
"source": self.source,
} }
@@ -445,6 +523,36 @@ def view_from_snapshot(snapshot: Any, now: Optional[float] = None,
) )
def view_from_socket_state(snapshot: Any, now: Optional[float] = None,
now_mono: Optional[float] = None) -> Optional[PluginRuntimeView]:
"""Judge the ``plugins`` section of a control-socket state snapshot by
the same rules as the cache copy; None when it has none (an older
display, or a snapshot too large to carry it), so the caller reads the
cache instead.
The display measured its render loop's heartbeat age when it answered
(``state.loop``); that is the heartbeat here, aged by the time since the
answer arrived. A live snapshot with a stalled loop is ``stalled``, and a
snapshot older than its ``stale_after`` (the publisher thread stopped)
is ``stale``, exactly as for the cache. The display answered, so its
process is alive: there is no pid check.
"""
from src.ipc.client import snapshot_loop_age # stdlib-only module
if not isinstance(snapshot, dict):
return None
state = snapshot.get("state")
plugins = state.get(STATE_SECTION) if isinstance(state, dict) else None
if not isinstance(plugins, dict):
return None
now_mono = time.monotonic() if now_mono is None else now_mono
beat_age = snapshot_loop_age(snapshot, now_mono=now_mono)
heartbeat = None
if beat_age is not None:
heartbeat = {"pid": plugins.get("pid"), "mono": now_mono - beat_age}
view = view_from_snapshot(plugins, now=now, heartbeat=heartbeat, now_mono=now_mono)
return replace(view, source="socket")
def read_plugin_runtime(cache_manager: Any, now: Optional[float] = None, def read_plugin_runtime(cache_manager: Any, now: Optional[float] = None,
heartbeat_path: Optional[str] = None) -> PluginRuntimeView: heartbeat_path: Optional[str] = None) -> PluginRuntimeView:
"""The display's latest snapshot, judged for staleness and against the """The display's latest snapshot, judged for staleness and against the
+108 -35
View File
@@ -8,7 +8,7 @@ Provides resource limits and performance monitoring.
import math import math
import time import time
import threading import threading
from typing import Dict, Optional, Any, Callable, cast from typing import Dict, Optional, Any, Callable, Set, cast
from dataclasses import dataclass, field, fields from dataclasses import dataclass, field, fields
from src.logging_config import get_logger from src.logging_config import get_logger
@@ -99,18 +99,33 @@ class ResourceMetrics:
last_update_time: float = field(default_factory=time.time) last_update_time: float = field(default_factory=time.time)
#: How often a plugin's metrics are written to the cache, in seconds. #: How often the metrics snapshot is written to the cache, in seconds.
#: #:
#: Persisting on every call meant a small file rewritten roughly nine times a #: Persisting on every call meant a small file rewritten roughly nine times a
#: minute per plugin. On a rig with fourteen active plugins that was ~126 #: minute per plugin. On a rig with fourteen active plugins that was ~126
#: writes a minute for metrics alone, and since each ~350-byte file costs a #: writes a minute for metrics alone, and since each ~350-byte file costs a
#: 4KB block plus an ext4 journal entry, it dominated the device's write #: 4KB block plus an ext4 journal entry, it dominated the device's write
#: volume -- on an SD card, which wears out. #: volume -- on an SD card, which wears out. Throttling each plugin's own
#: record to once per 30 s still left two writes a minute per plugin, so all
#: plugins now share one record (METRICS_SNAPSHOT_KEY), written at most once
#: a minute: one write a minute however many plugins there are.
#: #:
#: The in-memory copy stays authoritative and exact; only the cross-process #: The in-memory copy stays authoritative and exact; only the cross-process
#: snapshot the web UI reads is delayed, and telemetry up to half a minute old #: snapshot the web UI reads is delayed, and telemetry up to a minute old is
#: is still a fair description of a long-running plugin. #: still a fair description of a long-running plugin.
_METRICS_PERSIST_INTERVAL = 30.0 _METRICS_PERSIST_INTERVAL = 60.0
#: The one cache record holding every plugin's metrics:
#: ``{"schema": 1, "plugins": {plugin_id: <metrics record>}}``, each metrics
#: record shaped as the per-plugin ``plugin_metrics:<id>`` records were. Those
#: older records are still read for a plugin the snapshot does not have yet
#: (an upgrade, or a plugin that has not run since), never written.
METRICS_SNAPSHOT_KEY = "plugin_metrics_snapshot"
_METRICS_SNAPSHOT_SCHEMA = 1
#: A plugin with no call for this long is dropped from the snapshot -- what
#: the cache's 30-day default retention did to its own record before.
_METRICS_SNAPSHOT_ENTRY_MAX_AGE = 30 * 86400
class PluginResourceMonitor: class PluginResourceMonitor:
@@ -140,10 +155,15 @@ class PluginResourceMonitor:
self._metrics: Dict[str, ResourceMetrics] = {} self._metrics: Dict[str, ResourceMetrics] = {}
self._limits: Dict[str, ResourceLimits] = {} self._limits: Dict[str, ResourceLimits] = {}
self._bad_limits_warned: set = set() self._bad_limits_warned: set = set()
# When each plugin's metrics last reached the cache. Metrics change on # When the metrics snapshot last reached the cache (monotonic), None
# every call, so they cannot be de-duplicated the way health state can; # until it has. Metrics change on every call, so they cannot be
# they are rate-limited instead. See _METRICS_PERSIST_INTERVAL. # de-duplicated the way health state can; they are rate-limited
self._metrics_persisted_at: Dict[str, float] = {} # instead. See _METRICS_PERSIST_INTERVAL.
self._snapshot_persisted_at: Optional[float] = None
# Plugins whose metrics this process recorded since the last snapshot
# write: only their entries are overwritten, the rest are kept as
# found on disk.
self._metrics_dirty: Set[str] = set()
# Lock for thread-safe access # Lock for thread-safe access
self._lock = threading.Lock() self._lock = threading.Lock()
@@ -247,10 +267,14 @@ class PluginResourceMonitor:
with self._lock: with self._lock:
if force_reload or plugin_id not in self._metrics: if force_reload or plugin_id not in self._metrics:
# Try to load from cache # Try to load from cache
cache_key = self._get_metrics_key(plugin_id) memory_ttl = 0 if force_reload else None
cached = self._read_snapshot(memory_ttl).get(plugin_id)
if cached is None:
# Not in the snapshot: the per-plugin record an older
# version wrote, if there is one.
cached = self.cache_manager.get( cached = self.cache_manager.get(
cache_key, max_age=None, memory_ttl=0 if force_reload else None self._get_metrics_key(plugin_id), max_age=None,
) memory_ttl=memory_ttl)
if cached: if cached:
metrics = self._metrics_from_cache(plugin_id, cached) metrics = self._metrics_from_cache(plugin_id, cached)
else: else:
@@ -498,12 +522,70 @@ class PluginResourceMonitor:
summaries[plugin_id] = self.get_metrics_summary(plugin_id) summaries[plugin_id] = self.get_metrics_summary(plugin_id)
return summaries return summaries
def _persist_metrics(self, plugin_id: str, metrics: ResourceMetrics, def _read_snapshot(self, memory_ttl: Optional[int] = None) -> Dict[str, Any]:
force: bool = False) -> None: """The snapshot's per-plugin records, or {} if there is none usable.
"""Write a plugin's metrics to the cache, at most once per interval.
Caller must hold ``self._lock``. Caller must hold ``self._lock``.
""" """
cached = self.cache_manager.get(
METRICS_SNAPSHOT_KEY, max_age=None, memory_ttl=memory_ttl)
if not isinstance(cached, dict) or cached.get('schema') != _METRICS_SNAPSHOT_SCHEMA:
return {}
plugins = cached.get('plugins')
return plugins if isinstance(plugins, dict) else {}
@staticmethod
def _metrics_record(metrics: ResourceMetrics) -> Dict[str, Any]:
"""One plugin's entry in the snapshot."""
return {
'memory_mb': metrics.memory_mb,
'cpu_percent': metrics.cpu_percent,
'execution_time': metrics.execution_time,
'call_count': metrics.call_count,
'total_execution_time': metrics.total_execution_time,
'max_execution_time': metrics.max_execution_time,
'min_execution_time': (metrics.min_execution_time
if metrics.min_execution_time != float('inf')
else 0.0),
'last_update_time': metrics.last_update_time,
}
def _write_snapshot(self, drop: Optional[str] = None) -> None:
"""Write the snapshot: what is on disk, with this process's recorded
plugins updated and ``drop`` removed.
Starting from the disk copy rather than from memory keeps the entries
of plugins this process has not run -- disabled ones, which the web UI
still shows -- and a reset made from the other process.
Caller must hold ``self._lock``.
"""
plugins = dict(self._read_snapshot(memory_ttl=0))
if drop is not None:
plugins.pop(drop, None)
for plugin_id in self._metrics_dirty:
if plugin_id in self._metrics:
plugins[plugin_id] = self._metrics_record(self._metrics[plugin_id])
cutoff = time.time() - _METRICS_SNAPSHOT_ENTRY_MAX_AGE
for plugin_id, record in list(plugins.items()):
last = record.get('last_update_time') if isinstance(record, dict) else None
if isinstance(last, (int, float)) and last < cutoff:
del plugins[plugin_id]
self.cache_manager.set(METRICS_SNAPSHOT_KEY, {
'schema': _METRICS_SNAPSHOT_SCHEMA,
'plugins': plugins,
})
# Only once the write has landed, so a failed one is retried in full.
self._metrics_dirty.clear()
def _persist_metrics(self, plugin_id: str, metrics: ResourceMetrics,
force: bool = False) -> None:
"""Record that a plugin's metrics changed, and write the snapshot if
the last write is at least an interval old.
Caller must hold ``self._lock``.
"""
self._metrics_dirty.add(plugin_id)
# Monotonic, not wall clock: these devices have no RTC, so the clock # Monotonic, not wall clock: these devices have no RTC, so the clock
# jumps by however far off boot-time was the moment NTP first syncs. # jumps by however far off boot-time was the moment NTP first syncs.
# A forward jump would allow an early write, a backward one would # A forward jump would allow an early write, a backward one would
@@ -515,35 +597,26 @@ class PluginResourceMonitor:
# single run -- the throttle swallowed the very first snapshot, which # single run -- the throttle swallowed the very first snapshot, which
# is the one that matters most after a restart. # is the one that matters most after a restart.
now = time.monotonic() now = time.monotonic()
last_written = self._metrics_persisted_at.get(plugin_id) last_written = self._snapshot_persisted_at
if (not force and last_written is not None if (not force and last_written is not None
and now - last_written < _METRICS_PERSIST_INTERVAL): and now - last_written < _METRICS_PERSIST_INTERVAL):
return return
cache_key = self._get_metrics_key(plugin_id) self._write_snapshot()
self.cache_manager.set(cache_key, {
'memory_mb': metrics.memory_mb,
'cpu_percent': metrics.cpu_percent,
'execution_time': metrics.execution_time,
'call_count': metrics.call_count,
'total_execution_time': metrics.total_execution_time,
'max_execution_time': metrics.max_execution_time,
'min_execution_time': (metrics.min_execution_time
if metrics.min_execution_time != float('inf')
else 0.0),
'last_update_time': metrics.last_update_time,
})
# Only after the write lands. Marking it first would mean a failed # Only after the write lands. Marking it first would mean a failed
# set() bought the next interval's silence without leaving a snapshot. # set() bought the next interval's silence without leaving a snapshot.
self._metrics_persisted_at[plugin_id] = now self._snapshot_persisted_at = now
def reset_metrics(self, plugin_id: str) -> None: def reset_metrics(self, plugin_id: str) -> None:
"""Reset metrics for a plugin.""" """Reset metrics for a plugin."""
with self._lock: with self._lock:
if plugin_id in self._metrics: if plugin_id in self._metrics:
self._metrics[plugin_id] = ResourceMetrics() self._metrics[plugin_id] = ResourceMetrics()
cache_key = self._get_metrics_key(plugin_id) self._metrics_dirty.discard(plugin_id)
self.cache_manager.delete(cache_key) self._write_snapshot(drop=plugin_id)
# The record an older version wrote, so the reader's fallback
# cannot bring the old numbers back.
self.cache_manager.delete(self._get_metrics_key(plugin_id))
# Let the next call persist immediately rather than leaving the # Let the next call persist immediately rather than leaving the
# deleted key absent for the rest of the interval. # plugin absent from the snapshot for the rest of the interval.
self._metrics_persisted_at.pop(plugin_id, None) self._snapshot_persisted_at = None
+14
View File
@@ -344,12 +344,26 @@ class PluginStoreManager(_RegistryMixin, _InstallMixin, _UpdateMixin):
2. Fix permissions via os.chmod() then retry (works for same-owner files) 2. Fix permissions via os.chmod() then retry (works for same-owner files)
3. Use sudo rm -rf as last resort (works for root-owned __pycache__, etc.) 3. Use sudo rm -rf as last resort (works for root-owned __pycache__, etc.)
A symlink -- a dev plugin linked in by scripts/dev/dev_plugin_setup.sh
-- is removed as a link, before any of that: rmtree refuses one, and
stage 2 would walk through it and chmod the developer's checkout.
Args: Args:
path: Path to directory to remove path: Path to directory to remove
Returns: Returns:
True if directory was removed successfully, False otherwise True if directory was removed successfully, False otherwise
""" """
if path.is_symlink():
# Checked before exists(), which follows the link: a dangling one
# would read as already removed and be left behind.
try:
path.unlink()
return True
except OSError as e:
self.logger.error(f"Could not remove the symlink {path}: {e}")
return False
if not path.exists(): if not path.exists():
return True # Already removed return True # Already removed
+11 -8
View File
@@ -95,11 +95,13 @@ class StartupValidator:
def _validate_systemd_units(self) -> None: def _validate_systemd_units(self) -> None:
"""Warn when an installed unit has drifted from the repo's template. """Warn when an installed unit has drifted from the repo's template.
Nothing re-applies these after the first install. `git pull` -- which is Before updates refreshed units, nothing re-applied these after the
what the web UI's update button runs -- brings a new template into the first install: `git pull` brought a new template into the checkout,
checkout, but nothing copies it to /etc/systemd/system and nothing runs but nothing copied it to /etc/systemd/system, so the unit that
`systemctl daemon-reload`, so the unit that actually runs is whatever actually ran was whatever first_time_install.sh wrote on day one.
first_time_install.sh wrote on day one. Updates now install changed units through the root helper
ledmatrix-refresh-units (web_interface/unit_refresh.py) -- but only on
a device whose installer granted it, so this still catches the rest.
That makes every hardening added to a unit inert on existing installs. That makes every hardening added to a unit inert on existing installs.
Measured on one rig: the installed unit was thirteen days older than the Measured on one rig: the installed unit was thirteen days older than the
@@ -141,10 +143,11 @@ class StartupValidator:
if self._unit_body(expected) != self._unit_body(actual): if self._unit_body(expected) != self._unit_body(actual):
self.warnings.append( self.warnings.append(
f"{installed.name} differs from {template_rel}; the " f"{installed.name} differs from {template_rel}, so "
"installed unit is not refreshed by an update, so "
"settings added to the template are not in effect. " "settings added to the template are not in effect. "
"Re-run scripts/install/install_service.sh to apply them." "Updates apply them only once the installer has granted "
"ledmatrix-refresh-units: re-run "
"scripts/install/install_service.sh (or first_time_install.sh) to apply them."
) )
except OSError as e: except OSError as e:
self.logger.debug("Could not compare systemd units: %s", e) self.logger.debug("Could not compare systemd units: %s", e)
+5
View File
@@ -19,6 +19,11 @@ Type=simple
User=__USER__ User=__USER__
WorkingDirectory=__PROJECT_ROOT_DIR__ WorkingDirectory=__PROJECT_ROOT_DIR__
Environment=USE_THREADING=1 Environment=USE_THREADING=1
# Cap glibc's malloc arenas, as ledmatrix.service does: each allocating thread
# can get its own arena, up to 8 x CPU count (24 on a 3-core Pi), and a grown
# arena is never handed back to the OS. This threaded Flask process would hold
# that memory the same way. See ledmatrix.service for the measurement.
Environment=MALLOC_ARENA_MAX=2
ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/scripts/utils/start_web_conditionally.py ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/scripts/utils/start_web_conditionally.py
Restart=on-failure Restart=on-failure
RestartSec=10 RestartSec=10
+33
View File
@@ -328,6 +328,39 @@ def _hermetic_control_socket(monkeypatch):
monkeypatch.setenv(SOCKET_PATH_ENV, 'off') monkeypatch.setenv(SOCKET_PATH_ENV, 'off')
@pytest.fixture(autouse=True)
def _hermetic_unit_refresh(monkeypatch, tmp_path_factory):
"""Keep updates' systemd unit refresh off the host.
perform_core_update runs web_interface/unit_refresh.py after any update
that moves HEAD, and several tests run the real one against a test clone.
On a device -- or a machine where install_service.sh was tried out -- it
would compare the clone's templates with the real /etc/systemd/system and
run the real sudo helper. Point it at a folder that does not exist: no units
installed, nothing to do. The unit refresh tests pass their own.
"""
from web_interface import unit_refresh
monkeypatch.setattr(unit_refresh, 'SYSTEMD_DIR', str(tmp_path_factory.getbasetemp() / 'no-systemd'))
@pytest.fixture(autouse=True)
def _forget_settled_espn_chunks(monkeypatch):
"""src.common.espn_dates remembers past days process-wide; tests fake
different answers for the same dates, so none may inherit another's.
"Today" is also pinned to 2000-01-01, so no date a test uses counts as
settled unless the test says so (by pinning _utc_today itself). Without
that, a test asking for last month twice passes while that month is
recent and fails once it is three days old: the second ask is answered
from memory."""
from datetime import date
from src.common import espn_dates
monkeypatch.setattr(espn_dates, "_utc_today", lambda: date(2000, 1, 1))
espn_dates.clear_settled_chunk_cache()
yield
espn_dates.clear_settled_chunk_cache()
@pytest.fixture(autouse=True) @pytest.fixture(autouse=True)
def reset_logging(): def reset_logging():
"""Reset logging configuration before each test.""" """Reset logging configuration before each test."""
+4
View File
@@ -45,6 +45,10 @@ server has none.
|---|---|---| |---|---|---|
| `unit/test_list_filter.js` | no | `ListFilter` search/filter/sort/count/sticky, and the installed-plugins config **extracted verbatim** from `plugins_manager.js` so the test can't drift from it | | `unit/test_list_filter.js` | no | `ListFilter` search/filter/sort/count/sticky, and the installed-plugins config **extracted verbatim** from `plugins_manager.js` so the test can't drift from it |
| `unit/test_update_all.js` | no | `PluginInstallManager.updateAll` from `plugins/install_manager.js`: Check & Update All sends only plugin ids (never `starlark:` app entries), re-sends a request that got no HTTP answer (web service restarting) instead of skipping that plugin, never re-sends one that got any HTTP answer (the real `api_client.js` classifies a proxy 502 or a JSON error without `error_code` as `API_ERROR`), and counts a no-op update as already up to date in the summary. Also run by `test/web_interface/test_update_all_plugins.py` so CI covers it | | `unit/test_update_all.js` | no | `PluginInstallManager.updateAll` from `plugins/install_manager.js`: Check & Update All sends only plugin ids (never `starlark:` app entries), re-sends a request that got no HTTP answer (web service restarting) instead of skipping that plugin, never re-sends one that got any HTTP answer (the real `api_client.js` classifies a proxy 502 or a JSON error without `error_code` as `API_ERROR`), and counts a no-op update as already up to date in the summary. Also run by `test/web_interface/test_update_all_plugins.py` so CI covers it |
| `unit/test_store_install.js` | no | The store's Install button, with the whole of `plugins_manager.js` run by `plugins_manager_sandbox.js` (a vm context, fake DOM and API): a fresh install reloads the list, then enables the id the plugin was installed as -- the answer's `plugin_id`, else the installed entry the store entry matches (Weather installs as `ledmatrix-weather`); a Reinstall leaves the enabled state alone |
| `unit/test_install_polling.js` | no | How long Install waits for a queued install (sandbox): at least the server's 300 s dependency-install timeout; when it stops waiting it reloads the installed list and warns, rather than reporting a failure or enabling anything |
| `unit/test_store_categories.js` | no | The store's category filter (sandbox): the template ships only All Categories, the rest come from the store's plugins (one per category whatever its case), choosing one filters to it, and a swapped-in select is refilled from the cache keeping the choice |
| `unit/test_github_url_install.js` | no | Install Single Plugin (sandbox, the button as `plugins.html` ships it): no inline `onclick`, so a click or Enter sends exactly one `install-from-url` request and raises no error |
| `unit/test_render_cards.js` | no | `renderInstalledCards` markup, both empty states, and HTML-escaping of hostile plugin metadata | | `unit/test_render_cards.js` | no | `renderInstalledCards` markup, both empty states, and HTML-escaping of hostile plugin metadata |
| `unit/test_style_editor_element_keys.js` | no | `elementKeys()`/`styleRows()`/`positionRows()` from `widgets/style-editor.js`: every `customization.layout` entry gets exactly one row -- paired with its style element through core's `x-layout-key` (so `score` belongs to `score_text`, not a second row), or a position row of its own, leaves included -- since the widget claims the whole `layout` block from the generic fallback renderer | | `unit/test_style_editor_element_keys.js` | no | `elementKeys()`/`styleRows()`/`positionRows()` from `widgets/style-editor.js`: every `customization.layout` entry gets exactly one row -- paired with its style element through core's `x-layout-key` (so `score` belongs to `score_text`, not a second row), or a position row of its own, leaves included -- since the widget claims the whole `layout` block from the generic fallback renderer |
| `unit/test_style_editor_layout_leaf_columns.js` | no | `columnsFor()` from `widgets/style-editor.js`: a layout-only key whose own value is a leaf (no x/y sub-object, e.g. a `show_logo` toggle) gets a self-keyed column instead of a blank, uneditable row | | `unit/test_style_editor_layout_leaf_columns.js` | no | `columnsFor()` from `widgets/style-editor.js`: a layout-only key whose own value is a leaf (no x/y sub-object, e.g. a `show_logo` toggle) gets a self-keyed column instead of a blank, uneditable row |
+212
View File
@@ -0,0 +1,212 @@
// The whole of plugins_manager.js (and list_filter.js before it, as the page
// loads them), evaluated in a node vm context against a small fake DOM.
//
// For suites that drive the plugin manager's real flows -- install, polling,
// store filters, the GitHub-URL button -- rather than one function sliced
// out of the file. Nothing is mocked inside the script: only what the page
// gives it (document, fetch, timers, showNotification, LEDEscape).
//
// const sb = create({ route: (method, url, body) => ({ status, json }) });
// sb.el('plugin-store-grid'); // make an element exist by id
// sb.window.installPlugin('weather');
// await sb.until(() => sb.requests.some(r => r.url.includes('/toggle')));
//
// Timers ignore their delays and run on the next turn, so a poll loop that
// would take minutes in a browser finishes in milliseconds. The page is in
// readyState "loading" with no #installed-plugins-grid, so the script's own
// start-up does nothing until a suite asks for it (window.initPluginsPage()).
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const ledEscape = require('./led_escape');
const V3 = path.resolve(__dirname, '../../web_interface/static/v3');
const PLUGINS_HTML = path.resolve(__dirname, '../../web_interface/templates/v3/partials/plugins.html');
class FakeClassList {
constructor() { this.set = new Set(); }
add(...c) { c.forEach(x => this.set.add(x)); }
remove(...c) { c.forEach(x => this.set.delete(x)); }
contains(c) { return this.set.has(c); }
toggle(c, force) {
const on = force === undefined ? !this.set.has(c) : !!force;
if (on) this.set.add(c); else this.set.delete(c);
return on;
}
}
function create({ route } = {}) {
const elements = new Map();
const requests = [];
const toasts = [];
const errors = [];
const restartNotes = [];
class FakeElement {
constructor(id, tag = 'div', attributes = {}) {
this.id = id;
this.tagName = tag.toUpperCase();
this.attributes = { ...attributes };
this.listeners = {};
this.children = [];
this.classList = new FakeClassList();
this.style = { removeProperty() {} };
this.dataset = {};
this.value = '';
this.textContent = '';
this.disabled = false;
this.parentNode = null;
this._html = '';
}
get innerHTML() { return this._html; }
set innerHTML(v) { this._html = String(v); this.children = []; }
getAttribute(n) { return n in this.attributes ? this.attributes[n] : null; }
setAttribute(n, v) { this.attributes[n] = String(v); }
hasAttribute(n) { return n in this.attributes; }
removeAttribute(n) { delete this.attributes[n]; }
addEventListener(type, fn) { (this.listeners[type] = this.listeners[type] || []).push(fn); }
removeEventListener(type, fn) {
this.listeners[type] = (this.listeners[type] || []).filter(f => f !== fn);
}
appendChild(child) { this.children.push(child); child.parentNode = this; return child; }
querySelector() { return null; }
querySelectorAll() { return []; }
closest() { return null; }
cloneNode() {
const copy = new FakeElement(this.id, this.tagName, this.attributes);
copy._html = this._html;
copy.value = this.value;
return copy;
}
replaceChild(next, prev) {
next.parentNode = this;
prev.parentNode = null;
if (next.id) elements.set(next.id, next);
return prev;
}
replaceWith(next) { if (this.parentNode) this.parentNode.replaceChild(next, this); }
// A browser runs an inline on<type> attribute first (it was set before
// any listener was added), then the listeners, and an exception in one
// does not stop the next: it is reported, which is what `errors` holds.
dispatch(type, init = {}) {
const event = {
type, target: this, currentTarget: this, key: init.key,
defaultPrevented: false,
preventDefault() { this.defaultPrevented = true; },
stopPropagation() {}, stopImmediatePropagation() {},
};
const inline = this.getAttribute('on' + type);
const handlers = [];
if (inline !== null) {
handlers.push(vm.runInContext(`(function(event) {\n${inline}\n})`, ctx));
}
handlers.push(...(this.listeners[type] || []));
for (const h of handlers) {
try { h.call(this, event); } catch (e) { errors.push(e); }
}
return event;
}
click() { return this.dispatch('click'); }
}
function el(id, tag, attributes) {
if (!elements.has(id)) {
const parent = new FakeElement(null);
parent.appendChild(new FakeElement(id, tag, attributes));
elements.set(id, parent.children[0]);
}
return elements.get(id);
}
const timers = [];
const ctx = {
// Warnings are the script noting elements this fake page doesn't have.
console: { log: console.log.bind(console), error: console.error.bind(console),
warn: () => {}, info: () => {}, debug: () => {} },
debugLog: () => {},
addEventListener() {},
URL,
document: {
readyState: 'loading',
body: { addEventListener() {} },
getElementById: id => elements.get(id) || null,
querySelector: () => null,
querySelectorAll: () => [],
addEventListener() {},
dispatchEvent() { return true; },
createElement: tag => new FakeElement(null, tag),
},
CustomEvent: class { constructor(type, init) { this.type = type; this.detail = init && init.detail; } },
setTimeout: (fn, _ms, ...args) => { timers.push(setImmediate(() => fn(...args))); return timers.length; },
clearTimeout: () => {},
setInterval: () => 0,
clearInterval: () => {},
requestAnimationFrame: fn => setImmediate(fn),
getComputedStyle: () => ({ display: 'block' }),
scrollTo() {},
sessionStorage: { getItem: () => null, setItem() {}, removeItem() {} },
localStorage: { getItem: () => null, setItem() {}, removeItem() {} },
confirm: () => true,
alert: () => {},
showNotification: (message, type) => {
toasts.push({ message: String(message),
type: type && typeof type === 'object' ? type.type : type });
},
noteRestartRequired: (body) => { restartNotes.push(body); },
fetch: async (url, opts = {}) => {
const method = (opts.method || 'GET').toUpperCase();
let body = null;
try { body = opts.body ? JSON.parse(opts.body) : null; } catch (e) { body = opts.body; }
requests.push({ method, url: String(url), body });
const answer = (route && route(method, String(url), body)) || { status: 200, json: { status: 'success' } };
const status = answer.status || 200;
return { ok: status < 400, status, json: async () => answer.json };
},
};
ctx.window = ctx;
vm.createContext(ctx);
ledEscape.install(ctx);
for (const file of ['js/plugins/list_filter.js', 'plugins_manager.js']) {
vm.runInContext(fs.readFileSync(path.join(V3, file), 'utf8'), ctx, { filename: file });
}
// Resolves once cond() is true, letting timers and promises run between
// checks; rejects if it never is.
async function until(cond, label = 'condition', turns = 20000) {
for (let i = 0; i < turns; i++) {
if (cond()) return;
await new Promise(r => setImmediate(r));
}
throw new Error('timed out waiting for ' + label);
}
// Lets every pending timer and promise run.
async function settle(turns = 50) {
for (let i = 0; i < turns; i++) await new Promise(r => setImmediate(r));
}
return { window: ctx, el, FakeElement, requests, toasts, errors, restartNotes, until, settle };
}
// The attributes of the element with this id in partials/plugins.html, as
// the template ships them (no Jinja on the tags these suites read).
function templateAttributes(id) {
const html = fs.readFileSync(PLUGINS_HTML, 'utf8');
const at = html.indexOf(`id="${id}"`);
if (at < 0) throw new Error(`no element with id ${id} in plugins.html`);
const start = html.lastIndexOf('<', at);
let end = start, quote = null;
for (; end < html.length; end++) {
const ch = html[end];
if (quote) { if (ch === quote) quote = null; } else if (ch === '"' || ch === "'") quote = ch;
else if (ch === '>') break;
}
const tag = html.slice(start, end + 1);
const attrs = {};
const re = /([\w:-]+)\s*=\s*("([^"]*)"|'([^']*)')/g;
let m;
while ((m = re.exec(tag))) attrs[m[1]] = m[3] !== undefined ? m[3] : m[4];
return { tag: tag.match(/^<(\w+)/)[1], attrs, source: tag };
}
module.exports = { create, templateAttributes };
+6 -1
View File
@@ -20,7 +20,12 @@ const UNIT = ['unit/test_list_filter.js', 'unit/test_render_cards.js',
'unit/test_html_escaping.js', 'unit/test_style_editor_element_keys.js', 'unit/test_html_escaping.js', 'unit/test_style_editor_element_keys.js',
'unit/test_style_editor_layout_leaf_columns.js', 'unit/test_style_editor_layout_leaf_columns.js',
'unit/test_style_editor_layout_leaf_collision.js', 'unit/test_style_editor_layout_leaf_collision.js',
'unit/test_update_all.js', 'unit/test_inline_handler_escaping.js', 'unit/test_update_all.js',
'unit/test_store_install.js',
'unit/test_install_polling.js',
'unit/test_store_categories.js',
'unit/test_github_url_install.js',
'unit/test_inline_handler_escaping.js',
'unit/test_plugin_action_delegation.js', 'unit/test_file_upload_widget.js', 'unit/test_plugin_action_delegation.js', 'unit/test_file_upload_widget.js',
'unit/test_store_registry_fields.js', 'unit/test_restart_banner.js', 'unit/test_store_registry_fields.js', 'unit/test_restart_banner.js',
'unit/test_page_registry.js', 'unit/test_core_modules.js']; 'unit/test_page_registry.js', 'unit/test_core_modules.js'];
+93
View File
@@ -0,0 +1,93 @@
// Plugin Manager > Install from GitHub > Install Single Plugin: one click,
// one request, no errors.
//
// The Install button carried an inline onclick calling
// window.handleGitHubPluginInstall, and attachInstallButtonHandler also gave
// it a click listener that installs. Both ran on every click. The inline one
// threw a ReferenceError (it called isGithubUrl, which lives inside the
// plugin-manager IIFE, from outside it), so only the listener's request went
// out -- and fixing that scope alone would have sent every install twice.
// The button now has the listener only.
//
// Runs the whole of plugins_manager.js in the sandbox, with the button as
// partials/plugins.html ships it.
const { create, templateAttributes } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
const URL = 'https://github.com/someone/ledmatrix-demo';
function route(method, url) {
if (method === 'POST' && url === '/api/v3/plugins/install-from-url') {
return { json: { status: 'success', message: 'Plugin demo installed successfully', plugin_id: 'demo' } };
}
if (url.startsWith('/api/v3/plugins/installed')) return { json: { status: 'success', data: { plugins: [] } } };
return { json: { status: 'success' } };
}
function page() {
const sb = create({ route });
const button = templateAttributes('install-plugin-from-url');
sb.el('install-plugin-from-url', button.tag, button.attrs);
sb.el('github-plugin-url', 'input');
sb.el('github-plugin-status');
sb.el('plugin-branch-input', 'input');
return sb;
}
const installs = sb => sb.requests.filter(r => r.url === '/api/v3/plugins/install-from-url');
(async () => {
console.log('\nthe template');
{
const { attrs } = templateAttributes('install-plugin-from-url');
ok('the Install button has no inline onclick', !('onclick' in attrs), attrs.onclick);
}
console.log('\na click');
{
const sb = page();
sb.window.attachInstallButtonHandler();
// htmx:afterSettle runs it again on every swap; that must not add a handler.
sb.window.attachInstallButtonHandler();
sb.el('github-plugin-url').value = URL;
sb.window.document.getElementById('install-plugin-from-url').click();
await sb.settle();
ok('raises no error', sb.errors.length === 0, sb.errors.map(String));
ok('sends exactly one install request', installs(sb).length === 1, installs(sb));
ok('for the URL typed', installs(sb)[0] && installs(sb)[0].body.repo_url === URL, installs(sb));
ok('and reports the result', /Successfully installed: demo/.test(sb.el('github-plugin-status').innerHTML),
sb.el('github-plugin-status').innerHTML);
}
console.log('\nEnter in the URL field');
{
const sb = page();
sb.window.attachInstallButtonHandler();
const input = sb.el('github-plugin-url');
input.value = URL;
input.dispatch('keypress', { key: 'Enter' });
await sb.settle();
ok('raises no error', sb.errors.length === 0, sb.errors.map(String));
ok('sends exactly one install request', installs(sb).length === 1, installs(sb));
}
console.log('\na URL that is not GitHub');
{
const sb = page();
sb.window.attachInstallButtonHandler();
sb.el('github-plugin-url').value = 'https://example.com/x';
sb.window.document.getElementById('install-plugin-from-url').click();
await sb.settle();
ok('is refused without a request or an error',
installs(sb).length === 0 && sb.errors.length === 0 && /valid GitHub URL/.test(sb.el('github-plugin-status').innerHTML),
{ errors: sb.errors.map(String), status: sb.el('github-plugin-status').innerHTML });
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+89
View File
@@ -0,0 +1,89 @@
// How long the store's Install waits for a queued install, and what it says
// when it stops waiting.
//
// It polled the operation 60 times, a second apart, then reported "Install
// operation timed out" as an error and did nothing else. The server is
// allowed far longer: the plugin's dependency install alone may take 300 s
// (install_requirements_file in src/plugin_system/store_install.py), after
// a download that fetches the plugin one file at a time. So an install that
// went on to succeed was reported as failed, never enabled, and missing from
// the installed list until the page was reloaded.
//
// Runs the whole of plugins_manager.js in the sandbox; its timers ignore
// their delays, so each poll here stands for one second on a real page.
const { create } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
// The server's dependency-install timeout, in polls (one a second).
const DEPENDENCY_INSTALL_TIMEOUT_POLLS = 300;
function server(completesAfterPolls) {
const state = { polls: 0, installed: [] };
state.route = (method, url, body) => {
if (url.startsWith('/api/v3/plugins/installed')) {
return { json: { status: 'success', data: { plugins: state.installed.map(p => ({ ...p })) } } };
}
if (method === 'POST' && url === '/api/v3/plugins/install') {
return { json: { status: 'success', message: 'queued', data: { operation_id: 'op-1' } } };
}
if (url === '/api/v3/plugins/operation/op-1') {
state.polls++;
if (completesAfterPolls === null || state.polls < completesAfterPolls) {
return { json: { status: 'success', data: { status: 'running' } } };
}
state.installed = [{ id: 'clock-simple', name: 'Clock', enabled: false }];
return { json: { status: 'success', data: { status: 'completed',
result: { success: true, message: 'installed', plugin_id: 'clock-simple' } } } };
}
if (method === 'POST' && url === '/api/v3/plugins/toggle') {
return { json: { status: 'success', message: 'enabled' } };
}
return { json: { status: 'success' } };
};
return state;
}
(async () => {
console.log('\nan install that takes longer than a minute');
{
// 200 s: well inside what the server allows.
const srv = server(200);
const sb = create({ route: srv.route });
sb.window.installPlugin('clock-simple');
await sb.until(() => sb.toasts.some(t => /installed and enabled|enabling it failed|timed out|still/i.test(t.message)),
'the install to finish');
await sb.settle();
ok('is waited for until it completes', srv.polls === 200, srv.polls);
ok('is not reported as an error', !sb.toasts.some(t => t.type === 'error'), sb.toasts);
ok('and is enabled', sb.requests.some(r => r.url === '/api/v3/plugins/toggle' && r.body.plugin_id === 'clock-simple'),
sb.requests.filter(r => r.method === 'POST'));
}
console.log('\nan install that never reports back');
{
const srv = server(null);
const sb = create({ route: srv.route });
sb.window.installPlugin('clock-simple');
await sb.until(() => sb.toasts.length >= 3, 'the poller to give up');
await sb.settle();
ok(`is polled for at least the ${DEPENDENCY_INSTALL_TIMEOUT_POLLS} s dependency-install timeout`,
srv.polls >= DEPENDENCY_INSTALL_TIMEOUT_POLLS, srv.polls);
ok('...but not forever', srv.polls <= 1200, srv.polls);
const lastPoll = sb.requests.map(r => r.url).lastIndexOf('/api/v3/plugins/operation/op-1');
ok('then the installed list is reloaded, to show what actually happened',
sb.requests.slice(lastPoll + 1).some(r => r.url === '/api/v3/plugins/installed'),
sb.requests.slice(lastPoll + 1).map(r => r.url));
const last = sb.toasts[sb.toasts.length - 1];
ok('it says the install may still be running, as a warning, not a failure',
last && last.type === 'warning' && !/fail|timed out/i.test(last.message), sb.toasts);
ok('nothing is enabled on a guess', !sb.requests.some(r => r.url === '/api/v3/plugins/toggle'));
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+105
View File
@@ -0,0 +1,105 @@
// The Plugin Store's category filter offers the categories its plugins have.
//
// The template listed seven fixed categories. The registry uses about
// twenty (productivity, utility, transit, finance, ...), so roughly a third
// of the store could not be filtered to at all, and "Financial" missed the
// plugin filed under "finance". The options are now built from the store's
// plugins, as the Starlark section builds its own; the template ships only
// "All Categories".
//
// Runs the whole of plugins_manager.js in the sandbox.
const fs = require('fs');
const path = require('path');
const { create, templateAttributes } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
const STORE = [
{ id: 'nfl', name: 'NFL', category: 'sports' },
{ id: 'nba', name: 'NBA', category: 'Sports' },
{ id: 'todo', name: 'Todo', category: 'productivity' },
{ id: 'stocks', name: 'Stocks', category: 'finance' },
{ id: 'crypto', name: 'Crypto', category: 'financial' },
{ id: 'bus', name: 'Bus', category: 'transit' },
{ id: 'mystery', name: 'Mystery' },
];
function route(method, url) {
if (url.startsWith('/api/v3/plugins/store/list')) return { json: { status: 'success', data: { plugins: STORE } } };
if (url.startsWith('/api/v3/plugins/installed')) return { json: { status: 'success', data: { plugins: [] } } };
if (url.startsWith('/api/v3/plugins/store/github-status')) {
return { json: { status: 'success', data: { token_status: 'valid', authenticated: true, rate_limit: 5000 } } };
}
if (url.startsWith('/api/v3/plugins/saved-repositories')) {
return { json: { status: 'success', data: { repositories: [] } } };
}
if (url.startsWith('/api/v3/display/on-demand/status')) {
return { json: { status: 'success', data: { state: {}, service: {} } } };
}
return { json: { status: 'success' } };
}
const options = sel => sel.children.map(o => o.value);
const cardIds = sb => [...sb.el('plugin-store-grid').innerHTML.matchAll(/<h4[^>]*>([^<]*)<\/h4>/g)].map(m => m[1]);
(async () => {
console.log('\nthe template');
{
const html = fs.readFileSync(path.resolve(__dirname,
'../../../web_interface/templates/v3/partials/plugins.html'), 'utf8');
const start = html.indexOf('<select id="plugin-category"');
const block = html.slice(start, html.indexOf('</select>', start));
const shipped = [...block.matchAll(/<option value="([^"]*)"/g)].map(m => m[1]);
ok('ships only "All Categories"', JSON.stringify(shipped) === JSON.stringify(['']), shipped);
}
const sb = create({ route });
const attrs = templateAttributes('plugin-category').attrs;
const select = sb.el('plugin-category', 'select', attrs);
sb.el('plugin-store-grid');
sb.el('installed-plugins-grid');
sb.window.initPluginsPage();
await sb.until(() => sb.requests.some(r => r.url.startsWith('/api/v3/plugins/store/list')), 'the store list');
await sb.settle();
console.log('\noptions come from the store\'s plugins');
ok('"All Categories" is still the first choice', /<option value="">All Categories<\/option>/.test(select.innerHTML),
select.innerHTML);
ok('every category a plugin has is offered once, whatever its case',
JSON.stringify(options(select)) === JSON.stringify(['finance', 'financial', 'productivity', 'sports', 'transit']),
options(select));
ok('labels are capitalised',
(select.children.find(o => o.value === 'productivity') || {}).textContent === 'Productivity');
console.log('\nchoosing one filters to it');
select.value = 'productivity';
select.dispatch('change');
ok('productivity shows its plugin', JSON.stringify(cardIds(sb)) === JSON.stringify(['Todo']), cardIds(sb));
select.value = 'finance';
select.dispatch('change');
ok('finance is not lost to "financial"', JSON.stringify(cardIds(sb)) === JSON.stringify(['Stocks']), cardIds(sb));
select.value = 'sports';
select.dispatch('change');
ok('one option covers both spellings of sports',
JSON.stringify(cardIds(sb).sort()) === JSON.stringify(['NBA', 'NFL']), cardIds(sb));
console.log('\nthe partial is swapped back in (tab switch)');
{
// A fresh <select> from the template, the store list still cached.
const fresh = new sb.FakeElement('plugin-category', 'select', attrs);
select.parentNode.replaceChild(fresh, select);
sb.window.searchPluginStore(false);
await sb.settle();
ok('the new select is filled from the cache',
JSON.stringify(options(fresh)) === JSON.stringify(['finance', 'financial', 'productivity', 'sports', 'transit']),
options(fresh));
ok('keeping the chosen category', fresh.value === 'sports', fresh.value);
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+138
View File
@@ -0,0 +1,138 @@
// The store's Install button: which plugin it enables afterwards, and when.
//
// 1. Weather, Music, Stocks and Leaderboard are registry entries (`weather`)
// whose manifests declare another id (`ledmatrix-weather`). The plugin
// list, its config section and /plugins/toggle know them by that id, but
// the button enabled the registry id: /plugins/toggle answered 404
// "Plugin not found" and the plugin stayed disabled behind "installed,
// but enabling it failed". It now enables the id the install answer
// names (`plugin_id`), or, from an answer without one, the installed
// entry the store entry matches (its id, plugin_path name or aliases).
//
// 2. Reinstall (the same button on an installed plugin) enabled it too, so
// reinstalling a plugin the user had switched off switched it back on.
// Only a fresh install enables.
//
// Runs the whole of plugins_manager.js in the sandbox against a fake API.
const { create } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
const STORE = [
{ id: 'weather', name: 'Weather', category: 'weather', plugin_path: 'plugins/ledmatrix-weather',
aliases: ['ledmatrix-weather'] },
{ id: 'clock-simple', name: 'Clock', category: 'time', plugin_path: 'plugins/clock-simple', aliases: [] },
];
// A server with one install in flight. `queue` false answers the install
// directly; `names` false leaves plugin_id out of the answer (an older
// server); `installsAs` is the id the installed manifest declares.
function server({ installed = [], queue = true, names = true, installsAs }) {
const state = { installed: installed.map(p => ({ ...p })), polls: 0 };
const done = (id) => {
if (!state.installed.some(p => p.id === installsAs)) {
state.installed.push({ id: installsAs, name: id, enabled: false });
}
const result = { success: true, message: `Plugin ${id} installed successfully`, restart_required: false };
if (names) result.plugin_id = installsAs;
return result;
};
state.route = (method, url, body) => {
if (url.startsWith('/api/v3/plugins/store/list')) {
return { json: { status: 'success', data: { plugins: STORE } } };
}
if (url.startsWith('/api/v3/plugins/installed')) {
return { json: { status: 'success', data: { plugins: state.installed.map(p => ({ ...p })) } } };
}
if (method === 'POST' && url === '/api/v3/plugins/install') {
if (queue) return { json: { status: 'success', message: 'queued', data: { operation_id: 'op-1' } } };
return { json: { status: 'success', message: 'Plugin installed successfully', ...done(body.plugin_id) } };
}
if (url === '/api/v3/plugins/operation/op-1') {
state.polls++;
if (state.polls < 3) return { json: { status: 'success', data: { status: 'running' } } };
return { json: { status: 'success', data: { status: 'completed', result: done('weather') } } };
}
if (method === 'POST' && url === '/api/v3/plugins/toggle') {
const plugin = state.installed.find(p => p.id === body.plugin_id);
if (!plugin) return { status: 404, json: { status: 'error', message: 'Plugin not found' } };
plugin.enabled = body.enabled;
return { json: { status: 'success', message: `Plugin ${body.plugin_id} enabled successfully` } };
}
return { json: { status: 'success' } };
};
return state;
}
async function install(pluginId, opts) {
const srv = server(opts);
const sb = create({ route: srv.route });
sb.window.searchPluginStore();
await sb.until(() => sb.requests.some(r => r.url.startsWith('/api/v3/plugins/store/list')), 'store list');
await sb.window.pluginManager.loadInstalledPlugins(true);
await sb.settle();
sb.requests.length = 0;
sb.toasts.length = 0;
sb.window.installPlugin(pluginId);
await sb.until(() => sb.toasts.some(t => /installed and enabled|enabling it failed|reinstalled/.test(t.message)),
'the install to finish');
await sb.settle();
const toggles = sb.requests.filter(r => r.url === '/api/v3/plugins/toggle').map(r => r.body);
return { sb, srv, toggles };
}
(async () => {
console.log('\n1. a fresh install enables the id the plugin was installed as');
{
const { srv, toggles, sb } = await install('weather', { installsAs: 'ledmatrix-weather' });
ok('enables ledmatrix-weather, not the registry id',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'ledmatrix-weather', enabled: true }]), toggles);
ok('...which the server enabled', srv.installed.find(p => p.id === 'ledmatrix-weather').enabled === true, srv.installed);
ok('says so', sb.toasts.some(t => t.type === 'success' && /installed and enabled/.test(t.message)), sb.toasts);
ok('no "Plugin not found"', !sb.toasts.some(t => /not found|failed/.test(t.message)), sb.toasts);
const lastList = sb.requests.map(r => r.url).lastIndexOf('/api/v3/plugins/installed');
const toggleAt = sb.requests.findIndex(r => r.url === '/api/v3/plugins/toggle');
ok('the installed list is reloaded before enabling, so the new card is there to update',
lastList >= 0 && lastList < toggleAt, sb.requests.map(r => r.method + ' ' + r.url));
}
{
const { toggles } = await install('weather', { installsAs: 'ledmatrix-weather', names: false });
ok('an answer without plugin_id: the installed entry the store entry matches (its alias)',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'ledmatrix-weather', enabled: true }]), toggles);
}
{
const { toggles } = await install('weather', { installsAs: 'ledmatrix-weather', queue: false });
ok('without the operation queue, from the direct answer',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'ledmatrix-weather', enabled: true }]), toggles);
}
{
const { toggles } = await install('clock-simple', { installsAs: 'clock-simple', names: false });
ok('a plugin installed under its registry id is enabled by that id',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'clock-simple', enabled: true }]), toggles);
}
console.log('\n2. a reinstall leaves the plugin as the user had it');
{
const { srv, toggles, sb } = await install('weather', {
installsAs: 'ledmatrix-weather', installed: [{ id: 'ledmatrix-weather', name: 'Weather', enabled: false }],
});
ok('sends no toggle', toggles.length === 0, toggles);
ok('the plugin stays disabled', srv.installed.find(p => p.id === 'ledmatrix-weather').enabled === false);
ok('says it was reinstalled', sb.toasts.some(t => t.type === 'success' && /reinstalled/.test(t.message)), sb.toasts);
ok('and reloads the list',
sb.requests.some(r => r.url === '/api/v3/plugins/installed'), sb.requests.map(r => r.url));
}
{
const { toggles } = await install('weather', {
installsAs: 'ledmatrix-weather', installed: [{ id: 'ledmatrix-weather', name: 'Weather', enabled: true }],
});
ok('an enabled plugin is not toggled either', toggles.length === 0, toggles);
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+2 -1
View File
@@ -52,7 +52,8 @@ global.installedPlugins = [];
// eslint-disable-next-line no-eval // eslint-disable-next-line no-eval
eval([ eval([
'function escapeHtml(text) {', 'function escapeAttribute(text) {', 'function jsStringAttr(value) {', 'function escapeHtml(text) {', 'function escapeAttribute(text) {', 'function jsStringAttr(value) {',
'function isStorePluginInstalled(pluginIdOrPlugin) {', 'function renderPluginStore(plugins) {', 'function isStorePluginInstalled(pluginIdOrPlugin) {',
'function findInstalledStorePlugin(pluginIdOrPlugin) {', 'function renderPluginStore(plugins) {',
].map(extract).join('\n') + '\nglobal.renderPluginStore = renderPluginStore;' ].map(extract).join('\n') + '\nglobal.renderPluginStore = renderPluginStore;'
+ '\nglobal.isStorePluginInstalled = isStorePluginInstalled;'); + '\nglobal.isStorePluginInstalled = isStorePluginInstalled;');
+60 -3
View File
@@ -52,7 +52,7 @@ function fakeApi(behaviour = {}) {
}; };
} }
function setup(api, { stateList, windowList } = {}) { function setup(api, { stateList, windowList, pluginManager } = {}) {
global.window = { global.window = {
PluginAPI: api, PluginAPI: api,
installedPlugins: windowList, installedPlugins: windowList,
@@ -60,6 +60,7 @@ function setup(api, { stateList, windowList } = {}) {
installedPlugins: stateList, installedPlugins: stateList,
loadInstalledPlugins: async () => stateList, loadInstalledPlugins: async () => stateList,
}, },
pluginManager,
}; };
} }
@@ -92,12 +93,68 @@ const noSleep = { sleep: async () => {} };
ok('progress total counts only what is sent', ok('progress total counts only what is sent',
progress.length === EXPECTED.length && progress.every(([, n]) => n === EXPECTED.length), progress); progress.length === EXPECTED.length && progress.every(([, n]) => n === EXPECTED.length), progress);
} }
{
// A page without the plugin manager has no window.installedPlugins.
const api = fakeApi();
setup(api, { stateList: INSTALLED });
await Manager.updateAll(null, noSleep);
ok('the PluginStateManager list (no live list) is filtered the same way',
JSON.stringify(api.calls) === JSON.stringify(EXPECTED), api.calls);
}
console.log('\na second run sends the live list, not the first run\'s snapshot');
{
// Run 1 leaves PluginStateManager holding a, b, c. Then c is uninstalled
// and d installed: plugins_manager.js publishes that only as
// window.installedPlugins. Run 2 used to send a, b, c -- c failed as
// "plugin not found" and d, which had an update waiting, was skipped.
const api = fakeApi({
c: () => { throw { error_code: 'PLUGIN_UPDATE_FAILED', message: 'Plugin update failed: plugin not found' }; },
});
const stale = [{ id: 'a' }, { id: 'b' }, { id: 'c' }];
setup(api, { stateList: stale, windowList: [{ id: 'a' }, { id: 'b' }, { id: 'd' }] });
const results = await Manager.updateAll(null, noSleep);
ok('sends exactly what is installed now',
JSON.stringify(api.calls) === JSON.stringify(['a', 'b', 'd']), api.calls);
ok('...so nothing fails over an uninstalled plugin', results.every(r => r.success), results);
}
{ {
const api = fakeApi(); const api = fakeApi();
setup(api, { stateList: INSTALLED, windowList: [] }); setup(api, { stateList: INSTALLED, windowList: [] });
const results = await Manager.updateAll(null, noSleep);
ok('an empty live list means nothing is installed: nothing is sent',
api.calls.length === 0 && results.length === 0, api.calls);
}
console.log('\nthe end-of-run refresh redraws the installed grid');
{
// PluginStateManager's refresh replaced window.installedPlugins and
// nothing else: the cards kept "Update to vX" and the Updates badge
// kept its count. The plugin manager's load renders the grid.
const loads = [];
let stateLoads = 0;
const pluginManager = { loadInstalledPlugins: async (force) => { loads.push(force); } };
setup(fakeApi(), { stateList: INSTALLED, windowList: INSTALLED, pluginManager });
window.PluginStateManager.loadInstalledPlugins = async () => { stateLoads++; };
await Manager.updateAll(null, noSleep); await Manager.updateAll(null, noSleep);
ok('the PluginStateManager list is filtered the same way', ok('reloads through the plugin manager once, forced past its caches',
JSON.stringify(api.calls) === JSON.stringify(EXPECTED), api.calls); JSON.stringify(loads) === JSON.stringify([true]), loads);
ok('...instead of PluginStateManager', stateLoads === 0, stateLoads);
}
{
const pluginManager = { loadInstalledPlugins: async () => { throw new Error('offline'); } };
const answer = { status: 'success', data: { update_status: 'updated' }, restart_required: true };
setup(fakeApi({ 'ledmatrix-flights': () => answer }), { windowList: INSTALLED, pluginManager });
const warn = console.warn;
console.warn = () => {};
let results;
try {
results = await Manager.updateAll(null, noSleep);
} finally {
console.warn = warn;
}
ok('a failed plugin-manager reload still returns the results with their restart flag',
Array.isArray(results) && Manager.restartRequest(results) === answer);
} }
{ {
const api = fakeApi(); const api = fakeApi();
+41 -5
View File
@@ -145,13 +145,49 @@ class TestContextImageCache:
a = ctx.fit_image(img, (20, 20)) a = ctx.fit_image(img, (20, 20))
assert ctx.fit_image(img, (20, 20)) is a assert ctx.fit_image(img, (20, 20)) is a
def test_id_safety_pins_source(self, ctx): def test_id_keyed_entry_does_not_pin_source(self, ctx):
# id()-keyed entries must pin the source image so a recycled id import gc
# can't alias a dead image's cache entry. import weakref
img = _solid(10, 10) img = _solid(10, 10)
ctx.fit_image(img, (20, 20)) ctx.fit_image(img, (20, 20))
pinned = [entry[1] for entry in ctx._image_cache.values()] assert len(ctx._image_cache) == 1
assert img in pinned watch = weakref.ref(img)
del img
gc.collect()
assert watch() is None # the cache did not keep it alive
assert len(ctx._image_cache) == 0 # and its entry went with it
def test_fresh_image_each_frame_holds_nothing(self, ctx):
# draw_image(Image.open(path), box) every frame: the old pinning
# filled all 64 slots with dead-weight sources.
for _ in range(ctx._IMAGE_CACHE_MAX * 2):
ctx.fit_image(_solid(50, 50), (20, 20))
assert len(ctx._image_cache) == 0
def test_recycled_id_does_not_alias(self, ctx):
# A same-size image at a recycled address must not get the dead
# image's fit, even if the entry somehow outlived its source.
red = _solid(10, 10, (255, 0, 0, 255))
first = ctx.fit_image(red, (20, 20))
key = next(iter(ctx._image_cache))
entry = ctx._image_cache[key]
ctx._image_cache[key] = (entry[0], lambda: None) # source "gone"
again = ctx.fit_image(red, (20, 20))
assert again is not first # refit, not a stale hit
assert ctx.fit_image(red, (20, 20)) is again
def test_unweakrefable_source_is_pinned(self, ctx, monkeypatch):
import src.adaptive_layout as layout_mod
def no_weakref(*_args, **_kwargs):
raise TypeError("cannot create weak reference")
monkeypatch.setattr(layout_mod.weakref, "ref", no_weakref)
img = _solid(10, 10)
a = ctx.fit_image(img, (20, 20))
assert ctx.fit_image(img, (20, 20)) is a
assert next(iter(ctx._image_cache.values()))[1]() is img
def test_cache_key_entries_do_not_pin(self, ctx): def test_cache_key_entries_do_not_pin(self, ctx):
img = _solid(10, 10) img = _solid(10, 10)
@@ -0,0 +1,104 @@
"""POST /plugins/install says which id the plugin was installed as.
A registry entry can install under another id: `weather` (aliases
`ledmatrix-weather`) installs a directory whose manifest declares
`ledmatrix-weather`, and that is the id the plugin list, the plugin's config
section and /plugins/toggle know it by. The store's Install button enabled
the new plugin by the registry id, which /plugins/toggle answered with 404
"Plugin not found", so Weather, Music, Stocks and Leaderboard installed
disabled behind an "enabling it failed" warning.
The answer -- the queued operation's result, or the direct response --
carries `plugin_id`: the id the installed manifest declares, found the way
the store's update and uninstall find an install.
"""
import json
from unittest.mock import MagicMock
import pytest
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401
INSTALL = "/api/v3/plugins/install"
@pytest.fixture
def store(api_v3_module, tmp_path):
manager = api_v3_module.api_v3.plugin_store_manager
manager.install_plugin.return_value = True
manager.get_registry_info.return_value = None
manager._find_plugin_path.return_value = None
def installed_as(directory, manifest):
path = tmp_path / directory
path.mkdir()
(path / "manifest.json").write_text(json.dumps(manifest), encoding="utf-8")
manager._find_plugin_path.side_effect = (
lambda pid: path if pid == "weather" else None)
return path
manager.installed_as = installed_as
return manager
@pytest.fixture
def queued(api_v3_module):
queue = MagicMock()
def enqueue(operation_type, plugin_id, operation_callback=None):
queue.callback_result = operation_callback(MagicMock())
return "op-1"
queue.enqueue_operation.side_effect = enqueue
api_v3_module.api_v3.operation_queue = queue
return queue
class TestDirectInstall:
def test_an_aliased_entry_reports_the_manifest_id(self, api_v3_client, store):
store.installed_as("ledmatrix-weather", {"id": "ledmatrix-weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["status"] == "success"
assert body["plugin_id"] == "ledmatrix-weather"
store._find_plugin_path.assert_called_with("weather")
def test_an_entry_installed_under_its_own_id_reports_that(self, api_v3_client, store):
store.installed_as("weather", {"id": "weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["plugin_id"] == "weather"
def test_an_install_that_cannot_be_found_reports_the_requested_id(self, api_v3_client, store):
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["status"] == "success"
assert body["plugin_id"] == "weather"
def test_a_manifest_id_that_is_not_a_plain_name_is_not_passed_on(self, api_v3_client, store):
store.installed_as("ledmatrix-weather", {"id": "../elsewhere"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["plugin_id"] == "weather"
def test_an_unreadable_manifest_reports_the_requested_id(self, api_v3_client, store):
path = store.installed_as("ledmatrix-weather", {})
(path / "manifest.json").write_text("[not json", encoding="utf-8")
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["plugin_id"] == "weather"
def test_the_restart_fields_are_still_sent(self, api_v3_client, store):
store.installed_as("ledmatrix-weather", {"id": "ledmatrix-weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert "restart_required" in body
class TestQueuedInstall:
def test_the_operation_result_names_the_manifest_id(self, api_v3_client, store, queued):
store.installed_as("ledmatrix-weather", {"id": "ledmatrix-weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["data"]["operation_id"] == "op-1"
assert queued.callback_result["success"] is True
assert queued.callback_result["plugin_id"] == "ledmatrix-weather"
def test_an_install_that_cannot_be_found_names_the_requested_id(
self, api_v3_client, store, queued):
api_v3_client.post(INSTALL, json={"plugin_id": "weather"})
assert queued.callback_result["plugin_id"] == "weather"
@@ -0,0 +1,59 @@
"""GET /api/v3/plugins/installed carries each plugin's ``display_modes``.
The on-demand modal (plugins_manager.js) fills its Display Mode list from
``plugin.display_modes``, but the route never included the field, so every
plugin offered one option -- its own id -- under "This plugin exposes a
single display mode". The display turns that id into the plugin's first
mode, so a multi-mode plugin could only be started, and pinned, on that one.
The modes come from the plugin catalog (the manifests the web process
discovered), the same source /display/modes and on-demand/start use.
"""
from unittest.mock import MagicMock
import pytest
from test._api_v3_test_helpers import ( # noqa: F401 - fixtures
api_v3_client, api_v3_module,
)
@pytest.fixture
def installed(api_v3_module, api_v3_client, tmp_path):
def _get(declared_modes):
api = api_v3_module.api_v3
# The listing's own metadata says nothing about modes: what the
# route reports must come from the catalog.
info = {'id': 'football-scoreboard', 'name': 'Football', 'version': '1.0.0'}
api.plugin_catalog.plugins_dir = str(tmp_path) # no manifest on disk
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[info])
api.plugin_catalog.get_plugin_display_modes = MagicMock(return_value=declared_modes)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={})
response = api_v3_client.get('/api/v3/plugins/installed')
assert response.status_code == 200
plugins = [p for p in response.get_json()['data']['plugins']
if p['id'] == 'football-scoreboard']
assert len(plugins) == 1
api.plugin_catalog.get_plugin_display_modes.assert_any_call('football-scoreboard')
return plugins[0]
return _get
def test_every_declared_mode_is_listed_in_order(installed):
modes = ['nfl_live', 'nfl_recent', 'nfl_upcoming']
assert installed(modes)['display_modes'] == modes
def test_a_single_mode_plugin_lists_its_one_mode(installed):
assert installed(['clock-simple'])['display_modes'] == ['clock-simple']
def test_no_declared_modes_is_an_empty_list(installed):
# The modal falls back to the plugin id for an empty list.
assert installed([])['display_modes'] == []
def test_a_hand_edited_manifest_cannot_put_non_strings_in_the_list(installed):
assert installed(['nfl_live', 7, None, {'x': 1}])['display_modes'] == ['nfl_live']
+27
View File
@@ -13,6 +13,7 @@ The invariants that keep this change safe:
inline path exactly. inline path exactly.
""" """
import asyncio
import os import os
import sys import sys
import threading import threading
@@ -209,6 +210,32 @@ class TestFailurePaths:
assert pm.get_plugin_lock(plugin_id).acquire(blocking=False) is True assert pm.get_plugin_lock(plugin_id).acquire(blocking=False) is True
pm.get_plugin_lock(plugin_id).release() pm.get_plugin_lock(plugin_id).release()
@pytest.mark.parametrize("raised", [asyncio.CancelledError, SystemExit])
def test_update_raising_a_base_exception_still_releases_the_plugin(self, pm, raised):
"""asyncio.CancelledError and SystemExit derive from BaseException,
not Exception. Raised from update() on the worker, one skipped the
bookkeeping entirely: the plugin kept its lock and stayed RUNNING for
the life of the process -- never updated again, and every display()
skipped as busy."""
class CancellingPlugin(SlowPlugin):
def update(self):
self.update_calls += 1
raise raised()
plugin_id = _install(pm, CancellingPlugin())
pm.run_scheduled_updates()
deadline = time.monotonic() + 3
while pm.plugins[plugin_id].update_calls == 0 and time.monotonic() < deadline:
time.sleep(0.05)
time.sleep(0.2)
assert pm.get_plugin_lock(plugin_id).acquire(blocking=False) is True
pm.get_plugin_lock(plugin_id).release()
assert pm.state_manager.can_execute(plugin_id) is True
assert pm.plugin_last_update.get(plugin_id, 0) > 0
error = pm.state_manager.get_error_info(plugin_id)
assert error is not None and error["error_type"] == raised.__name__
def test_unloaded_while_queued_is_harmless(self, pm): def test_unloaded_while_queued_is_harmless(self, pm):
"""Exercise the public unload_plugin() lifecycle rather than """Exercise the public unload_plugin() lifecycle rather than
deleting pm.plugins directly: queue the target's update behind a deleting pm.plugins directly: queue the target's update behind a
+26
View File
@@ -517,6 +517,32 @@ class TestUpdateIsVerified:
h.updater.run() h.updater.run()
assert h.pending['dependency_failures'] == ['requirements.txt'] assert h.pending['dependency_failures'] == ['requirements.txt']
@pytest.mark.parametrize('unit_refresh, expected', [
({'status': 'refreshed', 'message': '', 'units': ['ledmatrix.service']}, True),
({'status': 'needs_reinstall', 'message': '', 'units': ['ledmatrix.service']}, False),
(None, False),
])
def test_the_health_check_learns_whether_the_update_installed_units(self, tmp_path, unit_refresh,
expected):
"""Its rollback restores the previous units only when this update replaced them."""
repo = Repo(tmp_path)
repo.publish()
h = Harness(tmp_path, repo, core_update=real_pull(repo.device, unit_refresh=unit_refresh))
h.updater.run()
assert h.pending['units_refreshed'] is expected
def test_a_health_check_that_never_starts_also_restores_the_units(self, tmp_path):
repo = Repo(tmp_path)
old = repo.head()
repo.publish()
refreshed = {'status': 'refreshed', 'message': '', 'units': ['ledmatrix.service']}
h = Harness(tmp_path, repo, pickup=False,
core_update=real_pull(repo.device, unit_refresh=refreshed))
h.updater.run()
assert repo.head() == old
assert [a for a in h.sudo if a[-1] == '--restore'] == [
['sudo', '-n', '/usr/local/sbin/ledmatrix-refresh-units', '--restore']]
def test_a_health_check_that_never_starts_means_the_update_is_undone(self, tmp_path): def test_a_health_check_that_never_starts_means_the_update_is_undone(self, tmp_path):
repo = Repo(tmp_path) repo = Repo(tmp_path)
old = repo.head() old = repo.head()
+48
View File
@@ -80,6 +80,8 @@ class FakeHost:
self.nrestarts = 0 self.nrestarts = 0
self.heartbeat = heartbeat self.heartbeat = heartbeat
self.display_started_at = -1000.0 # the pre-update display, long running self.display_started_at = -1000.0 # the pre-update display, long running
self.unit_restores = [] # (argv, commit checked out, restarts so far)
self.restore_ok = True
def broken(self, kind): def broken(self, kind):
if self.running_head is None: if self.running_head is None:
@@ -103,6 +105,10 @@ class FakeHost:
if self.pip: if self.pip:
return self.pip(args, self) return self.pip(args, self)
return done(args, rc=0 if self.pip_ok else 1) return done(args, rc=0 if self.pip_ok else 1)
if args[:3] == ['sudo', '-n', av.REFRESH_UNITS_PATH]:
self.unit_restores.append((list(args), git(self.repo, 'rev-parse', 'HEAD'),
len(self.restarts)))
return done(args, rc=0 if self.restore_ok else 1)
if args[:4] == ['sudo', '-n', 'systemctl', 'restart']: if args[:4] == ['sudo', '-n', 'systemctl', 'restart']:
if self.restart_failures: if self.restart_failures:
self.restart_failures -= 1 self.restart_failures -= 1
@@ -405,3 +411,45 @@ def test_the_heartbeat_location_and_freshness_match_the_display():
# window it has to stay healthy for. # window it has to stay healthy for.
assert av.HEARTBEAT_FRESH_SECONDS + av.POLL_SECONDS < av.STABLE_SECONDS assert av.HEARTBEAT_FRESH_SECONDS + av.POLL_SECONDS < av.STABLE_SECONDS
assert av.HEARTBEAT_FRESH_SECONDS > display_watchdog.BEAT_INTERVAL_SECONDS * 2 assert av.HEARTBEAT_FRESH_SECONDS > display_watchdog.BEAT_INTERVAL_SECONDS * 2
# -- systemd units the update installed ------------------------------------------
def test_a_rollback_restores_the_units_the_update_installed(tmp_path):
"""The update installed new units (web_interface/unit_refresh.py); the
rollback puts the old ones back before restarting onto the old code."""
code, result, host, head, old, new = check(tmp_path, 'display_down', units_refreshed=True)
assert result['status'] == 'rolled_back' and head == old
assert len(host.unit_restores) == 1
argv, commit, restarts_before = host.unit_restores[0]
assert argv == ['sudo', '-n', av.REFRESH_UNITS_PATH, '--restore']
assert commit == old, 'restored after the code was rolled back'
assert restarts_before == 2, 'restored before the services restart onto the old code'
assert host.restarts[-2:] == [('ledmatrix.service', old), ('ledmatrix-web.service', old)]
assert result['detail'] is None
@pytest.mark.parametrize('pending', [{}, {'units_refreshed': False}])
def test_a_rollback_leaves_units_alone_when_the_update_did_not_change_them(tmp_path, pending):
# {} is what an updater from before this change writes.
code, result, host, head, old, new = check(tmp_path, 'display_down', **pending)
assert result['status'] == 'rolled_back' and host.unit_restores == []
def test_a_healthy_update_keeps_its_new_units(tmp_path):
code, result, host, head, old, new = check(tmp_path, units_refreshed=True)
assert result['status'] == 'success' and host.unit_restores == []
def test_a_failed_unit_restore_is_reported_but_the_rollback_stands(tmp_path):
repo, old, new = updated_repo(tmp_path)
av.write_pending(av.pending_path(repo), {'status': 'pending', 'old_head': old, 'new_head': new,
'display_was_active': True, 'dependency_failures': [],
'units_refreshed': True})
host = FakeHost(repo, new, 'display_down')
host.restore_ok = False
host.verifier().verify()
result = av.read_pending(av.pending_path(repo))
assert result['status'] == 'rolled_back'
assert git(repo, 'rev-parse', 'HEAD') == old
assert 'install_service.sh' in result['detail']
+18
View File
@@ -200,6 +200,24 @@ class TestGetOdds:
manager.get_odds('football', 'nfl', '401') # hit manager.get_odds('football', 'nfl', '401') # hit
assert [r for r in caplog.records if r.levelno == logging.INFO] == [] assert [r for r in caplog.records if r.levelno == logging.INFO] == []
def test_debug_off_does_not_serialize_the_response(
self, manager, mock_get, caplog):
# json.dumps(indent=2) of every odds body ran even with DEBUG off.
with caplog.at_level(logging.INFO, logger=manager.logger.name), \
patch('src.base_odds_manager.json.dumps') as dumps:
assert manager.get_odds('football', 'nfl', '401') == FULL_EXTRACTED
dumps.assert_not_called()
def test_debug_on_still_logs_the_raw_response(
self, manager, mock_get, caplog):
with caplog.at_level(logging.DEBUG, logger=manager.logger.name):
manager.get_odds('football', 'nfl', '401')
messages = [r.getMessage() for r in caplog.records]
assert any(m.startswith('Received raw odds data from ESPN: {')
for m in messages), messages
assert any(m.startswith('Returning extracted odds data: {')
for m in messages), messages
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# _extract_espn_data # _extract_espn_data
+99
View File
@@ -0,0 +1,99 @@
"""A cache key too long to be a filename still gets a cache file.
The calendar plugin's key joins every calendar id the user picked; on a real
install it passed 300 bytes, and since ext4 caps a filename at 255 every write
failed with ENAMETOOLONG -- logged as "permission denied", every update.
"""
import logging
import os
from unittest.mock import patch
from src.cache.disk_cache import DiskCache, _MAX_KEY_FILENAME_BYTES, _filename_stem
from src.cache_manager import CacheManager
# The shape of the key that failed on hdpi, ids anonymised.
CALENDAR_KEY = (
"calendar_events_someone@example.com_en.usa#holiday@group.v.calendar.google.com_"
"family13997378751670666433@group.calendar.google.com_ncaaf_-m-07kbp5_"
"%47eorgia+%42ulldogs+football#sports@group.v.calendar.google.com_nfl_-m-07l24_"
"%54ampa+%42ay+%42uccaneers#sports@group.v.calendar.google.com_primary"
)
# ext4/xfs/btrfs NAME_MAX; set()'s temp file adds 15 bytes to the stem.
NAME_MAX = 255
TEMP_OVERHEAD = len(".") + len(".json") + len(".") + 8
def test_the_real_key_was_too_long_to_write():
assert len((CALENDAR_KEY + ".json").encode()) > NAME_MAX - 10
def test_a_long_key_round_trips(tmp_path):
cache = DiskCache(str(tmp_path))
cache.set(CALENDAR_KEY, {"events": [1, 2, 3]})
assert cache.get(CALENDAR_KEY, max_age=None) == {"events": [1, 2, 3]}
path = cache.get_cache_path(CALENDAR_KEY)
assert os.path.isfile(path)
stem = os.path.basename(path)[:-len(".json")]
assert len(stem.encode()) + TEMP_OVERHEAD <= NAME_MAX
def test_short_keys_keep_their_filename(tmp_path):
cache = DiskCache(str(tmp_path))
exactly = "k" * _MAX_KEY_FILENAME_BYTES
assert cache.get_cache_path("weather_current") == str(tmp_path / "weather_current.json")
assert cache.get_cache_path(exactly) == str(tmp_path / f"{exactly}.json")
assert cache.get_cache_path(exactly + "k") != str(tmp_path / f"{exactly}k.json")
def test_long_keys_sharing_a_prefix_stay_apart(tmp_path):
cache = DiskCache(str(tmp_path))
first, second = CALENDAR_KEY + "_a", CALENDAR_KEY + "_b"
cache.set(first, {"which": "a"})
cache.set(second, {"which": "b"})
assert cache.get_cache_path(first) != cache.get_cache_path(second)
assert cache.get(first, max_age=None) == {"which": "a"}
assert cache.get(second, max_age=None) == {"which": "b"}
def test_the_prefix_never_splits_a_character():
key = "news_" + "é" * 300 # two bytes each, so the cut lands mid-character
stem = _filename_stem(key)
assert stem.startswith("news_é")
assert len(stem.encode("utf-8")) <= _MAX_KEY_FILENAME_BYTES
stem.encode("utf-8").decode("utf-8") # well-formed
def test_a_stem_listed_by_the_web_ui_deletes_the_same_file(tmp_path):
with patch('src.cache_manager.CacheManager._get_writable_cache_dir', return_value=str(tmp_path)):
manager = CacheManager()
try:
manager.save_cache(CALENDAR_KEY, {"events": []})
listed = [entry["key"] for entry in manager.list_cache_files()]
assert len(listed) == 1
manager.clear_cache(listed[0])
assert [n for n in os.listdir(tmp_path) if n.endswith(".json")] == []
finally:
manager.stop_cleanup_thread()
def test_a_failed_write_names_the_real_error(tmp_path, monkeypatch, caplog):
blocker = tmp_path / "a-file"
blocker.write_text("")
# No writable fallback either, so set() gives up and says why.
monkeypatch.setattr(os.path, "expanduser", lambda _p: str(blocker / "home"))
cache = DiskCache(str(tmp_path / "missing"))
with caplog.at_level(logging.WARNING):
cache.set("weather_current", {"t": 1})
gave_up = [r.getMessage() for r in caplog.records if "Could not write cache" in r.getMessage()]
assert len(gave_up) == 1
assert "permission denied" not in gave_up[0]
assert os.strerror(2) in gave_up[0] # ENOENT: the directory does not exist
+98
View File
@@ -0,0 +1,98 @@
"""CacheManager builds its ConfigManager on first use, not in __init__.
Every CacheManager built a ConfigManager and loaded the whole config for a
cache strategy that no longer reads it. The attribute stays public -- the
sports plugins resolve the global timezone through
``cache_manager.config_manager`` -- so it is now built on first access.
"""
from unittest.mock import MagicMock, patch
import pytest
import src.config_manager as config_manager_module
from src.cache_manager import CacheManager
@pytest.fixture
def built(monkeypatch):
"""Count ConfigManager constructions and load_config calls."""
made = []
class CountingConfigManager:
def __init__(self):
made.append(self)
self.loads = 0
def load_config(self):
self.loads += 1
return {}
monkeypatch.setattr(config_manager_module, "ConfigManager", CountingConfigManager)
return made
@pytest.fixture
def manager(tmp_path):
with patch('src.cache_manager.CacheManager._get_writable_cache_dir',
return_value=str(tmp_path)):
cm = CacheManager()
cm.stop_cleanup_thread()
return cm
def test_construction_does_not_load_the_config(built, manager):
manager.set("k", {"v": 1})
assert manager.get("k") == {"v": 1}
assert manager.get_cache_strategy("sports_live")["max_age"] > 0
assert built == []
def test_first_access_builds_and_loads_it_once(built, manager):
first = manager.config_manager
assert manager.config_manager is first
assert getattr(manager, "config_manager", None) is first
assert len(built) == 1 and first.loads == 1
def test_assignment_still_wins(built, manager):
replacement = MagicMock()
manager.config_manager = replacement
assert manager.config_manager is replacement
assert built == []
def test_an_unimportable_config_manager_is_none(manager, monkeypatch):
import builtins
real_import = builtins.__import__
def refuse(name, *args, **kwargs):
if name == "src.config_manager":
raise ImportError("no config manager here")
return real_import(name, *args, **kwargs)
monkeypatch.setattr(builtins, "__import__", refuse)
assert manager.config_manager is None
def test_a_failed_load_is_retried_on_the_next_access(manager, monkeypatch):
attempts = []
class Flaky:
def load_config(self):
attempts.append(1)
if len(attempts) == 1:
raise RuntimeError("config.json unreadable")
return {}
monkeypatch.setattr(config_manager_module, "ConfigManager", Flaky)
with pytest.raises(RuntimeError):
manager.config_manager
assert isinstance(manager.config_manager, Flaky)
assert len(attempts) == 2
def test_a_manager_made_without_init_still_answers(built):
bare = CacheManager.__new__(CacheManager)
bare.logger = MagicMock()
assert bare.config_manager is built[0]
+95
View File
@@ -0,0 +1,95 @@
"""The memory tier never serves a record older than the reader asked for.
A record read from disk went into the memory tier timed from the read, not
from when it was written, so get(max_age=300) could hand out data up to twice
that old: after a restart, after the memory sweep, or in a second process that
loaded a record once and kept serving it.
"""
import time
from unittest.mock import patch
import pytest
from src.cache_manager import CacheManager
class Clock:
def __init__(self, now):
self.now = now
def __call__(self):
return self.now
@pytest.fixture
def clock(monkeypatch):
fake = Clock(1_800_000_000.0)
monkeypatch.setattr(time, "time", fake)
return fake
def _manager(path):
# No disk sweep: it judges files by their real mtime against the fake
# clock and would delete them as months old.
with patch('src.cache_manager.CacheManager._get_writable_cache_dir',
return_value=str(path)), \
patch('src.cache_manager.CacheManager.start_cleanup_thread'):
return CacheManager()
def test_a_record_loaded_late_expires_on_its_own_timestamp(tmp_path, clock):
writer = _manager(tmp_path)
writer.set("weather_current", {"t": 1})
reader = _manager(tmp_path) # a restart, or the other process
clock.now += 250
assert reader.get("weather_current", max_age=300) == {"t": 1}
clock.now += 100 # the data is 350 s old; it sat in memory for 100 s
assert reader.get("weather_current", max_age=300) is None
def test_a_stored_ttl_bounds_the_memory_copy_too(tmp_path, clock):
writer = _manager(tmp_path)
writer.set("odds_espn_football_nfl_401", {"spread": 6.5}, ttl=60)
reader = _manager(tmp_path)
clock.now += 55
assert reader.get("odds_espn_football_nfl_401", max_age=3600) == {"spread": 6.5}
clock.now += 60
assert reader.get("odds_espn_football_nfl_401", max_age=3600) is None
def test_a_stale_memory_copy_gives_way_to_a_newer_write_on_disk(tmp_path, clock):
writer = _manager(tmp_path)
reader = _manager(tmp_path)
writer.set("stocks_AAPL", {"price": 1})
assert reader.get("stocks_AAPL", max_age=300) == {"price": 1}
clock.now += 280
writer.set("stocks_AAPL", {"price": 2})
clock.now += 40 # reader's copy: 40 s in memory, 320 s old
assert reader.get("stocks_AAPL", max_age=300) == {"price": 2}
def test_fresh_records_are_still_served_from_memory(tmp_path, clock):
manager = _manager(tmp_path)
manager.set("news_NFL", {"items": []})
clock.now += 100
with patch.object(manager._disk_cache_component, "get") as disk_get:
assert manager.get("news_NFL", max_age=300) == {"items": []}
disk_get.assert_not_called()
def test_max_age_none_and_records_without_a_timestamp_never_expire(tmp_path, clock):
manager = _manager(tmp_path)
manager.set("plugin_health_x", {"ok": True})
manager.save_cache("raw_record", {"no": "timestamp"})
clock.now += 10 ** 6
assert manager.get("plugin_health_x", max_age=None) == {"ok": True}
assert manager.get_cached_data("raw_record", max_age=None) == {"no": "timestamp"}
+21
View File
@@ -116,3 +116,24 @@ def test_response_json_prefers_orjson_and_falls_back():
assert json_body.response_json(response) == payload assert json_body.response_json(response) == payload
# A response object without bytes content (a test double) still works. # A response object without bytes content (a test double) still works.
assert json_body.response_json(SimpleNamespace(json=lambda: payload)) == payload assert json_body.response_json(SimpleNamespace(json=lambda: payload)) == payload
@pytest.mark.parametrize("body", [b"not json", b"", b'{"a": NaN}', b"\xef\xbb\xbf{}"])
def test_response_json_raises_and_returns_what_requests_does(body):
# The core fetch paths (api_helper, base_odds_manager, logo_downloader,
# dynamic_team_resolver) catch requests' JSONDecodeError on a bad body, so
# response_json must raise exactly that, and parse whatever requests
# parses (NaN, which orjson rejects) to the same value.
import requests
response = requests.models.Response()
response._content = body
response.status_code = 200
response.headers["Content-Type"] = "application/json"
try:
expected = response.json()
except requests.exceptions.JSONDecodeError:
with pytest.raises(requests.exceptions.JSONDecodeError):
json_body.response_json(response)
else:
assert json.dumps(json_body.response_json(response)) == json.dumps(expected)
+242
View File
@@ -0,0 +1,242 @@
"""An unchanged CacheManager.set() does not rewrite the file, and the skip
never changes how old the record looks.
DiskCache.set already skipped a payload identical to the last one it wrote,
but CacheManager.set stamps every record with time.time(), so for set() the
payload was never identical and every unchanged re-save was a full rewrite on
the SD card. The digest now leaves the header timestamp out, and the newer
timestamp lives in the file's mtime instead ("UNCHANGED RE-SAVES" in
src/cache/disk_cache.py). These tests pin both halves: the write is skipped,
and every reader still ages the record from its newest save -- across a
restart, and with a bound on what a foreign mtime can claim.
"""
import os
import shutil
import tempfile
import time
from unittest.mock import patch
import pytest
from src.cache import disk_cache as disk_cache_module
from src.cache.disk_cache import (
DiskCache,
_MAX_TIMESTAMP_LIFT,
_effective_timestamp,
)
from src.cache_manager import CacheManager
@pytest.fixture
def writes(monkeypatch):
"""Count real writes: every atomic write starts with mkstemp."""
calls = []
real = tempfile.mkstemp
def counting(*args, **kwargs):
calls.append(kwargs.get("prefix"))
return real(*args, **kwargs)
monkeypatch.setattr(disk_cache_module.tempfile, "mkstemp", counting)
return calls
@pytest.fixture
def clock(monkeypatch):
"""time.time() for the cache modules, advanced by hand."""
now = [time.time()]
fake = type("FakeTime", (), {"time": staticmethod(lambda: now[0])})
monkeypatch.setattr(disk_cache_module, "time", fake)
import src.cache_manager as cache_manager_module
monkeypatch.setattr(cache_manager_module, "time", fake)
return now
@pytest.fixture
def cm(tmp_path):
with patch('src.cache_manager.CacheManager._get_writable_cache_dir',
return_value=str(tmp_path)):
manager = CacheManager()
manager.stop_cleanup_thread()
yield manager
def _record(ts, data=None, ttl=None):
"""A record laid out the way CacheManager.set writes it."""
rec = {"timestamp": ts}
if ttl is not None:
rec["ttl"] = ttl
rec["data"] = data if data is not None else {"games": [1, 2, 3]}
return rec
class TestTheWriteIsSkipped:
def test_repeated_identical_set_writes_once(self, cm, writes):
path = cm._get_cache_path("scores")
for _ in range(20):
cm.set("scores", {"games": [1, 2, 3]}, ttl=60)
assert len(writes) == 1
# The data is the same and the file is the same file.
assert cm.get("scores", max_age=60, memory_ttl=0) == {"games": [1, 2, 3]}
assert os.stat(path).st_nlink == 1
def test_the_file_is_not_replaced(self, tmp_path, clock):
disk = DiskCache(str(tmp_path))
disk.set("k", _record(clock[0]))
before = os.stat(disk.get_cache_path("k"))
clock[0] += 30
disk.set("k", _record(clock[0]))
after = os.stat(disk.get_cache_path("k"))
assert after.st_ino == before.st_ino
assert after.st_mtime == pytest.approx(clock[0], abs=1e-3)
def test_changed_data_rewrites(self, cm, writes):
cm.set("scores", {"games": [1]})
cm.set("scores", {"games": [2]})
assert len(writes) == 2
assert cm.get("scores", memory_ttl=0) == {"games": [2]}
def test_a_changed_ttl_rewrites(self, cm, writes):
cm.set("scores", {"games": [1]}, ttl=60)
cm.set("scores", {"games": [1]}, ttl=600)
cm.set("scores", {"games": [1]})
assert len(writes) == 3
def test_a_timestamp_going_backwards_rewrites(self, tmp_path, writes, clock):
disk = DiskCache(str(tmp_path))
disk.set("k", _record(clock[0]))
disk.set("k", _record(clock[0] - 100))
assert len(writes) == 2
assert disk.get("k", max_age=None)["timestamp"] == pytest.approx(clock[0] - 100)
def test_unchanged_data_is_still_rewritten_once_the_lift_runs_out(
self, tmp_path, writes, clock):
disk = DiskCache(str(tmp_path))
start = clock[0]
disk.set("k", _record(start))
clock[0] = start + _MAX_TIMESTAMP_LIFT - 1
disk.set("k", _record(clock[0]))
assert len(writes) == 1
clock[0] = start + _MAX_TIMESTAMP_LIFT + 1
disk.set("k", _record(clock[0]))
assert len(writes) == 2
# ...and the embedded timestamp caught up.
with open(disk.get_cache_path("k"), "rb") as f:
assert disk_cache_module._loads(f.read())["timestamp"] == clock[0]
def test_another_writer_replacing_the_file_forces_a_rewrite(self, tmp_path, writes):
mine, theirs = DiskCache(str(tmp_path)), DiskCache(str(tmp_path))
now = time.time()
mine.set("k", _record(now, {"v": "mine"}))
theirs.set("k", _record(now + 1, {"v": "theirs"}))
mine.set("k", _record(now + 2, {"v": "mine"}))
assert len(writes) == 3
assert mine.get("k", max_age=None)["data"] == {"v": "mine"}
def test_a_record_without_a_timestamp_still_skips(self, tmp_path, writes):
disk = DiskCache(str(tmp_path))
disk.set("k", {"plain": True})
disk.set("k", {"plain": True})
assert len(writes) == 1
class TestAgeAfterSkippedWrites:
def test_a_skipped_save_keeps_the_record_fresh(self, tmp_path, clock):
disk = DiskCache(str(tmp_path))
disk.set("k", _record(clock[0], ttl=60))
for _ in range(10): # ten minutes of unchanged 50 s re-saves
clock[0] += 50
disk.set("k", _record(clock[0], ttl=60))
clock[0] += 50
record = disk.get("k", max_age=300)
assert record is not None
# Handed back as a rewrite would have left it.
assert record["timestamp"] == pytest.approx(clock[0] - 50, abs=1e-3)
def test_and_it_expires_on_time_once_the_saves_stop(self, tmp_path, clock, monkeypatch):
disk = DiskCache(str(tmp_path))
disk.set("k", _record(clock[0]))
clock[0] += 200
disk.set("k", _record(clock[0]))
clock[0] += 59
assert disk.get("k", max_age=60) is not None
clock[0] += 2
parses = []
real = disk_cache_module._loads
monkeypatch.setattr(disk_cache_module, "_loads",
lambda raw: parses.append(1) or real(raw))
assert disk.get("k", max_age=60) is None
assert parses == [] # still decided from the header
def test_the_ttl_is_honoured_the_same_way(self, tmp_path, clock):
disk = DiskCache(str(tmp_path))
disk.set("k", _record(clock[0], ttl=30))
clock[0] += 100
disk.set("k", _record(clock[0], ttl=30))
clock[0] += 20
assert disk.get("k", max_age=5) is not None # ttl wins, 20 < 30
clock[0] += 20
assert disk.get("k", max_age=3600) is None # 40 > 30
def test_a_restart_sees_the_newest_save(self, tmp_path, clock, writes):
disk = DiskCache(str(tmp_path))
disk.set("k", _record(clock[0]))
clock[0] += 250
disk.set("k", _record(clock[0]))
restarted = DiskCache(str(tmp_path)) # empty digest map
assert restarted.get("k", max_age=60) is not None
# It rewrites once (it cannot know what is on disk), then skips.
clock[0] += 10
restarted.set("k", _record(clock[0]))
clock[0] += 10
restarted.set("k", _record(clock[0]))
assert len(writes) == 2
def test_the_cache_manager_reads_it_across_processes(self, cm, tmp_path, clock):
cm.set("display_state", {"mode": "clock"})
clock[0] += 100
cm.set("display_state", {"mode": "clock"})
# The web interface: its own manager, memory tier bypassed.
with patch('src.cache_manager.CacheManager._get_writable_cache_dir',
return_value=str(tmp_path)):
web = CacheManager()
web.stop_cleanup_thread()
clock[0] += 60
assert web.get("display_state", max_age=120, memory_ttl=0) == {"mode": "clock"}
clock[0] += 70
assert web.get("display_state", max_age=120, memory_ttl=0) is None
def test_retention_sees_the_newest_save(self, cm, clock):
cm.set("odds_x", {"line": 1})
path = cm._get_cache_path("odds_x")
clock[0] += 3000
cm.set("odds_x", {"line": 1})
assert os.path.getmtime(path) == pytest.approx(clock[0], abs=1e-3)
class TestFreshnessCannotBeBorrowed:
def test_a_record_written_with_an_old_timestamp_reads_old(self, tmp_path):
disk = DiskCache(str(tmp_path))
old = time.time() - 600
disk.set("k", _record(old))
assert os.path.getmtime(disk.get_cache_path("k")) == pytest.approx(old, abs=1e-3)
assert disk.get("k", max_age=300) is None
def test_a_copy_without_mtime_is_bounded(self, tmp_path):
disk = DiskCache(str(tmp_path))
day_old = time.time() - 86400
disk.set("k", _record(day_old))
copy_dir = tmp_path / "restored"
copy_dir.mkdir()
shutil.copyfile(disk.get_cache_path("k"), copy_dir / "k.json") # mtime = now
restored = DiskCache(str(copy_dir))
assert restored.get("k", max_age=300) is None
record = restored.get("k", max_age=None)
assert record["timestamp"] == pytest.approx(day_old + _MAX_TIMESTAMP_LIFT)
def test_effective_timestamp(self):
assert _effective_timestamp(1000.0, None) == 1000.0
assert _effective_timestamp(1000.0, 900.0) == 1000.0 # mtime older
assert _effective_timestamp(1000.0, 1500.0) == 1500.0 # a skipped save
assert _effective_timestamp(1000.0, 10 ** 9) == 1000.0 + _MAX_TIMESTAMP_LIFT
+199
View File
@@ -0,0 +1,199 @@
"""ConfigService notifies subscribers outside its lock, in order.
Subscribers ran while _load_config held the service's lock. The display's
per-plugin subscriber is PluginManager.apply_config_change, which waits up to
PLUGIN_LOCK_TIMEOUT (5 s) for a busy plugin. The same save that toggles a
plugin's ``enabled`` flags a reconcile, and the render thread runs it: its
get_config() -- and the unsubscribe() of a plugin it disables -- waited behind
every slow callback, freezing the panel for up to 5 s per busy plugin.
What callers could rely on before still holds: one reload's notifications
finish before the next reload's start, and a callback unsubscribe() removed is
not running, and will not run, once unsubscribe() returns.
"""
import itertools
import json
import os
import threading
import time
import pytest
from src.config_manager import ConfigManager
from src.config_service import ConfigService
SLOW = 2.0 # how long a blocked callback waits before giving up
@pytest.fixture
def service(tmp_path):
config_path = tmp_path / "config.json"
config_path.write_text(json.dumps({"display": {"brightness": 50},
"weather": {"enabled": True}}),
encoding="utf-8")
manager = ConfigManager(str(config_path), str(tmp_path / "config_secrets.json"))
manager.template_path = str(tmp_path / "no-template.json")
svc = ConfigService(manager, enable_hot_reload=False)
yield svc, config_path
svc.shutdown()
_saves = itertools.count(1)
def _save(config_path, **sections):
config = json.loads(config_path.read_text(encoding="utf-8"))
config.update(sections)
config_path.write_text(json.dumps(config), encoding="utf-8")
# ConfigManager re-reads only when (mtime, size) moves. Two quick saves of
# the same size can share an mtime tick (about 16 ms on Windows), so step
# it forward explicitly.
st = config_path.stat()
os.utime(config_path, ns=(st.st_atime_ns, st.st_mtime_ns + next(_saves) * 50_000_000))
def _reload_in_background(svc):
thread = threading.Thread(target=svc._load_config, daemon=True)
thread.start()
return thread
def test_get_config_does_not_wait_for_a_slow_subscriber(service):
svc, config_path = service
entered, release = threading.Event(), threading.Event()
def slow(_old, _new):
entered.set()
release.wait(SLOW)
svc.subscribe(slow, plugin_id="weather")
_save(config_path, weather={"enabled": False})
reload = _reload_in_background(svc)
assert entered.wait(SLOW)
start = time.monotonic()
config = svc.get_config()
waited = time.monotonic() - start
release.set()
reload.join(SLOW)
assert waited < 0.5
# Swapped before anyone was told: a subscriber that reads it sees the new one.
assert config["weather"]["enabled"] is False
def test_unsubscribing_another_callback_does_not_wait(service):
svc, config_path = service
entered, release = threading.Event(), threading.Event()
def slow(_old, _new):
entered.set()
release.wait(SLOW)
def other(_old, _new):
pass
svc.subscribe(slow, plugin_id="weather")
svc.subscribe(other, plugin_id="clock")
_save(config_path, weather={"enabled": False})
reload = _reload_in_background(svc)
assert entered.wait(SLOW)
start = time.monotonic()
svc.unsubscribe(other, plugin_id="clock")
waited = time.monotonic() - start
release.set()
reload.join(SLOW)
assert waited < 0.5
def test_a_callback_unsubscribed_mid_notification_is_not_called(service):
svc, config_path = service
entered, release = threading.Event(), threading.Event()
called = []
def slow_global(_old, _new): # global subscribers are notified first
entered.set()
release.wait(SLOW)
def weather(_old, _new):
called.append("weather")
svc.subscribe(slow_global)
svc.subscribe(weather, plugin_id="weather")
_save(config_path, weather={"enabled": False})
reload = _reload_in_background(svc)
assert entered.wait(SLOW)
svc.unsubscribe(weather, plugin_id="weather")
release.set()
reload.join(SLOW)
assert called == []
def test_unsubscribe_waits_for_its_own_callback_to_return(service):
svc, config_path = service
entered, release = threading.Event(), threading.Event()
returned = threading.Event()
def slow(_old, _new):
entered.set()
release.wait(SLOW)
returned.set()
svc.subscribe(slow, plugin_id="weather")
_save(config_path, weather={"enabled": False})
reload = _reload_in_background(svc)
assert entered.wait(SLOW)
threading.Timer(0.2, release.set).start()
svc.unsubscribe(slow, plugin_id="weather")
assert returned.is_set()
reload.join(SLOW)
def test_a_callback_may_read_config_and_unsubscribe_itself(service):
svc, config_path = service
seen = []
def once(_old, _new):
seen.append(svc.get_config()["weather"]["enabled"])
svc.unsubscribe(once, plugin_id="weather")
svc.subscribe(once, plugin_id="weather")
_save(config_path, weather={"enabled": False})
reload = _reload_in_background(svc)
reload.join(SLOW)
assert not reload.is_alive()
assert seen == [False]
def test_two_reloads_notify_in_order(service):
svc, config_path = service
entered, release = threading.Event(), threading.Event()
seen = []
def record(old, new):
seen.append((old["brightness"], new["brightness"]))
if len(seen) == 1:
entered.set()
release.wait(SLOW)
svc.subscribe(record, plugin_id="display")
_save(config_path, display={"brightness": 60})
first = _reload_in_background(svc)
assert entered.wait(SLOW)
_save(config_path, display={"brightness": 100})
second = _reload_in_background(svc)
time.sleep(0.2)
release.set()
first.join(SLOW)
second.join(SLOW)
assert seen == [(50, 60), (60, 100)]
+112
View File
@@ -0,0 +1,112 @@
"""A plugin duration that is not a number must not stop the display.
Several plugins return their ``display_duration`` setting as it is in
config.json (``return self.config.get('display_duration', 15.0)``), so a
value saved as ``"20"`` or ``null`` -- from the raw config editor, or by
hand -- reached run() as a string or None. _resolve_durations then compared
it with 0, the TypeError went past every handler in the loop, and the
display service exited; systemd restarted it into the same screen and the
same crash.
"""
import logging
import math
import os
from unittest.mock import MagicMock
os.environ.setdefault("EMULATOR", "true")
import pytest
from src.display_controller import DisplayController
from test._run_loop_harness import FakePlugin, RunLoopHarness
def _controller(plugin_modes):
dc = object.__new__(DisplayController)
dc.config = {}
dc.plugin_modes = plugin_modes
return dc
def _plugin(duration, plugin_id='clock-simple'):
plugin = MagicMock()
plugin.plugin_id = plugin_id
plugin.get_display_duration.return_value = duration
return plugin
class TestPluginDurationIsCoerced:
@pytest.mark.parametrize('value, expected', [
('20', 20.0), (' 7.5 ', 7.5), (12, 12.0), (12.5, 12.5)])
def test_numbers_and_numeric_strings_are_used(self, value, expected):
dc = _controller({'clock': _plugin(value)})
duration = dc._get_display_duration('clock')
assert duration == expected and isinstance(duration, float)
@pytest.mark.parametrize('value', [
None, '', 'twenty', True, False, float('nan'), float('inf'), 'inf',
[20], {'seconds': 20}])
def test_anything_but_a_finite_number_gets_the_default(self, value):
dc = _controller({'clock': _plugin(value)})
assert dc._get_display_duration('clock') == 30
@pytest.mark.parametrize('value', [0, -5, '-5', '0'])
def test_a_number_not_above_zero_still_gets_the_15s_rule(self, value):
"""Unchanged: _resolve_durations turns it into 15 s, with its warning."""
plugin = _plugin(value)
dc = _controller({'clock': plugin})
base = dc._get_display_duration('clock')
assert dc._resolve_durations(plugin, 'clock', base, False)[1] == 15.0
def test_a_raising_get_display_duration_gets_the_default(self):
plugin = _plugin(None)
plugin.get_display_duration.side_effect = KeyError('display_duration')
assert _controller({'clock': plugin})._get_display_duration('clock') == 30
def test_the_result_feeds_resolve_durations(self):
"""The two calls run() makes back to back, for one screen."""
plugin = _plugin('bad')
dc = _controller({'clock': plugin})
base = dc._get_display_duration('clock')
assert dc._resolve_durations(plugin, 'clock', base, False) == (30, 30)
def test_logged_once_per_plugin(self, caplog):
dc = _controller({'clock': _plugin('twenty'),
'clock_big': _plugin('twenty'),
'calendar': _plugin(None, plugin_id='calendar')})
# clock_big is a second mode of the same plugin.
dc.plugin_modes['clock_big'].plugin_id = 'clock-simple'
with caplog.at_level(logging.WARNING, logger='src.display_controller'):
for _ in range(3):
for mode in ('clock', 'clock_big', 'calendar'):
dc._get_display_duration(mode)
warnings = [r for r in caplog.records if 'display duration' in r.getMessage()]
assert len(warnings) == 2
assert {'clock-simple', 'calendar'} == {
next(p for p in ('clock-simple', 'calendar') if p in r.getMessage())
for r in warnings}
def test_a_good_value_after_a_bad_one_is_used(self):
plugin = _plugin(None)
dc = _controller({'clock': plugin})
assert dc._get_display_duration('clock') == 30
plugin.get_display_duration.return_value = 45
assert dc._get_display_duration('clock') == 45.0
class TestRunLoopSurvives:
"""Through the real run() on the harness's fake clock."""
@pytest.mark.parametrize('duration, shown_for', [('20', 20.0), (None, 30.0),
('twenty', 30.0)])
def test_the_screen_runs_and_the_rotation_goes_on(self, tmp_path, duration, shown_for):
harness = RunLoopHarness(tmp_path, horizon=120)
harness.add_plugin(FakePlugin("weather", ["weather"], duration=30))
harness.add_plugin(FakePlugin("clock-simple", ["clock"], duration=duration))
# Before the fix run() returned at t=30, when the clock came up, and
# the harness raised "run() returned ... before the horizon".
rows = harness.run()["screens"]
clock = next(row for row in rows if row[1] == "clock")
assert math.isclose(clock[2], shown_for, abs_tol=1.0)
assert [row[1] for row in rows][:3] == ["weather", "clock", "weather"]
+12
View File
@@ -118,6 +118,18 @@ class TestDisplayManagerResourceManagement:
dm.matrix.Clear.assert_called() dm.matrix.Clear.assert_called()
def test_cleanup_takes_the_gc_monitor_out_of_gc_callbacks(
self, test_config, mock_rgb_matrix):
"""The service stops (SIGTERM -> run()'s finally -> cleanup()) with
its collection timer unregistered, not left for interpreter teardown."""
import gc
with patch.dict('os.environ', {'EMULATOR': 'false'}):
dm = DisplayManager(test_config)
monitor = dm.frame_timing.gc_monitor
assert monitor in gc.callbacks
dm.cleanup()
assert monitor not in gc.callbacks
class TestDisplayManagerDoubleSided: class TestDisplayManagerDoubleSided:
"""Double-sided mode: render once at logical size, tile across the chain.""" """Double-sided mode: render once at logical size, tile across the chain."""
+100
View File
@@ -39,6 +39,14 @@ def forget_rejected_ranges(monkeypatch):
monkeypatch.setattr(espn_dates, "_ranges_rejected_until", 0.0) monkeypatch.setattr(espn_dates, "_ranges_rejected_until", 0.0)
@pytest.fixture(autouse=True)
def nothing_is_settled_yet(monkeypatch):
"""Pin "today" before every date these tests use, so the settled-chunk
memory stays out of tests that are not about it whatever the real date.
TestSettledChunkCache moves it forward."""
monkeypatch.setattr(espn_dates, "_utc_today", lambda: date(2000, 1, 1))
class FakeResponse: class FakeResponse:
def __init__(self, status_code=200, payload=None): def __init__(self, status_code=200, payload=None):
self.status_code = status_code self.status_code = status_code
@@ -489,3 +497,95 @@ class TestConcurrency:
assert live["peak"] <= espn_dates.ESPN_CHUNK_WORKERS assert live["peak"] <= espn_dates.ESPN_CHUNK_WORKERS
assert live["peak"] > 1, "chunks should actually overlap" assert live["peak"] > 1, "chunks should actually overlap"
class TestSettledChunkCache:
"""Days that ended three or more days ago are fetched once a day, not hourly.
The scoreboards re-fetch a 22-day window every hour; on hdpi (2026-10-02)
the 12 settled days were 68% of that window's bytes.
"""
TODAY = date(2026, 10, 2)
# The scoreboards' default window on that day: 14 back, 7 ahead.
WINDOW = "20260918-20261009"
@pytest.fixture(autouse=True)
def frozen_today(self, monkeypatch):
monkeypatch.setattr(espn_dates, "_utc_today", lambda: self.TODAY)
def _events(self):
days = [(9, d) for d in range(18, 31)] + [(10, d) for d in range(1, 10)]
return {"2026%02d%02d" % (m, d): [{"id": f"{m}-{d}"}] for m, d in days}
def test_the_second_refresh_only_asks_for_unsettled_days(self):
session = FakeSession(self._events())
first = fetch_espn_scoreboard(session, URL, params={"dates": self.WINDOW})
session.calls.clear()
second = fetch_espn_scoreboard(session, URL, params={"dates": self.WINDOW})
asked = sorted(call["dates"] for call in session.calls)
# Sep 29 is the last settled day (today minus three).
assert asked == ["20260930"] + ["202610%02d" % d for d in range(1, 10)]
assert second["events"] == first["events"] # same events, same order
assert len(second["events"]) == 22
def test_a_hit_is_a_fresh_copy(self):
session = FakeSession(self._events())
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
hit = fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
hit["events"][0]["id"] = "mutated"
again = fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
assert again["events"][0]["id"] == "9-18"
def test_other_params_are_part_of_the_key(self):
session = FakeSession(self._events())
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW, "groups": 80})
session.calls.clear()
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
assert len(session.calls) == 22 # a different question, nothing reused
def test_entries_expire_after_a_day(self, monkeypatch):
clock = [1000.0]
monkeypatch.setattr(espn_dates.time, "monotonic", lambda: clock[0])
session = FakeSession(self._events())
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
clock[0] += espn_dates.SETTLED_CHUNK_TTL_SECONDS + 1
session.calls.clear()
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
assert len(session.calls) == 22
def test_failed_chunks_are_not_remembered(self):
session = FakeSession(self._events(), fail_chunks={"20260920"})
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
session.fail_chunks.clear()
session.calls.clear()
data = fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
assert "20260920" in [call["dates"] for call in session.calls]
assert len(data["events"]) == 22
def test_a_capped_month_is_not_remembered_but_its_days_are(self):
full = [{"id": f"x{i}"} for i in range(ESPN_MAX_LIMIT)]
session = FakeSession({"202608": full, "20260801": [{"id": "d1"}]})
fetch_espn_date_chunks(session, URL, params={"dates": "20260801-20260831"})
session.calls.clear()
data = fetch_espn_date_chunks(session, URL, params={"dates": "20260801-20260831"})
assert [call["dates"] for call in session.calls] == ["202608"]
assert [event["id"] for event in data["events"]] == ["d1"]
def test_memory_is_bounded(self, monkeypatch):
monkeypatch.setattr(espn_dates, "SETTLED_CACHE_MAX_ENTRIES", 5)
session = FakeSession(self._events())
fetch_espn_date_chunks(session, URL, params={"dates": self.WINDOW})
assert len(espn_dates._settled_chunks) == 5
assert espn_dates._settled_bytes == sum(
len(blob) for _, blob in espn_dates._settled_chunks.values())
def test_single_day_requests_are_untouched(self):
# The live path asks for today (or one day) as a plain request; that
# never goes through chunks or the memory.
session = FakeSession(self._events())
for _ in range(2):
fetch_espn_scoreboard(session, URL, params={"dates": "20260918"})
assert len(session.calls) == 2
+5
View File
@@ -13,6 +13,7 @@ No network: sessions are fakes, and the fetch service is a fresh one per test.
import json import json
import logging import logging
import os
import threading import threading
import time import time
from datetime import date, datetime from datetime import date, datetime
@@ -321,6 +322,10 @@ class TestWithARealCacheManager:
record["timestamp"] = time.time() - seconds record["timestamp"] = time.time() - seconds
with open(path, "w", encoding="utf-8") as fh: with open(path, "w", encoding="utf-8") as fh:
json.dump(record, fh) json.dump(record, fh)
# Rewriting the file moves its mtime to now, and the disk cache takes a
# record's age from the newer of its timestamp and its mtime (an
# unchanged re-save only touches the file), so age the mtime as well.
os.utime(path, (record["timestamp"], record["timestamp"]))
cm._memory_cache_component.clear() cm._memory_cache_component.clear()
def test_a_writers_long_ttl_does_not_outlast_the_readers(self, cm, service): def test_a_writers_long_ttl_does_not_outlast_the_readers(self, cm, service):
@@ -76,3 +76,44 @@ def test_only_the_latest_image_is_adopted():
dc._adopt_follower_scroll_image(rp) dc._adopt_follower_scroll_image(rp)
assert helper.cached_image is latest assert helper.cached_image is latest
assert helper.total_scroll_width == 400 assert helper.total_scroll_width == 400
def test_a_follower_frame_does_not_build_a_deferred_strip_image():
"""Every follower frame asks whether the strip is there. A strip the
follower's own rebuild deferred (append_content) must be answered from
the helper's bookkeeping, not by building and keeping its PIL image."""
import numpy as np
from src.common.scroll_helper import ScrollHelper
width, height = 64, 16
helper = ScrollHelper(width, height)
helper.append_content([Image.new("RGB", (300, height), (0, 200, 0))])
helper.append_content([Image.new("RGB", (300, height), (0, 0, 200))])
assert helper.__dict__.get("_cached_image") is None
assert helper.has_strip()
dc = object.__new__(DisplayController)
dc.config = {"sync": {}}
dc.vegas_coordinator = SimpleNamespace(
render_pipeline=SimpleNamespace(scroll_helper=helper))
dc.display_manager = MagicMock(width=width)
dc.sync_manager = MagicMock()
dc.sync_manager.get_latest_scroll_x.return_value = 200
dc._follower_incoming_image = deque(maxlen=1)
dc._follower_dr_last_t = None
dc._follower_local_x = 200.0
dc._follower_pending_new_image = False
dc._follower_last_frame = None
dc._follower_deadline = None
dc._scroll_speed = 0
for _ in range(3):
dc._run_follower_frame()
# The frame was cut from the array...
frame = np.asarray(dc._follower_last_frame)
assert frame.shape == (height, width, 3)
assert frame.any()
# ...and the strip is still held once.
assert helper.__dict__.get("_cached_image") is None
+69 -1
View File
@@ -626,7 +626,7 @@ def test_render_bench_strip_lights_a_real_share_of_pixels():
def _collection(monitor, monkeypatch, start, took, generation=2): def _collection(monitor, monkeypatch, start, took, generation=2):
"""One collection of ``took`` seconds, as gc.callbacks would report it.""" """One collection of ``took`` seconds, as gc.callbacks would report it."""
clock = iter([start, start + took]) clock = iter([start, start + took])
monkeypatch.setattr(frame_timing.time, "perf_counter", lambda: next(clock)) monkeypatch.setattr(monitor, "_clock", lambda: next(clock))
monitor("start", {"generation": generation}) monitor("start", {"generation": generation})
monitor("stop", {"generation": generation, "collected": 0, "uncollectable": 0}) monitor("stop", {"generation": generation, "collected": 0, "uncollectable": 0})
monkeypatch.undo() monkeypatch.undo()
@@ -739,3 +739,71 @@ def test_installing_the_gc_monitor_twice_installs_it_once():
first = frame_timing.install_gc_monitor() first = frame_timing.install_gc_monitor()
assert frame_timing.install_gc_monitor() is first assert frame_timing.install_gc_monitor() is first
assert sum(1 for cb in gc.callbacks if cb is first) == 1 assert sum(1 for cb in gc.callbacks if cb is first) == 1
def test_uninstalling_the_gc_monitor_removes_it_and_can_repeat():
import gc
first = frame_timing.install_gc_monitor()
frame_timing.uninstall_gc_monitor()
frame_timing.uninstall_gc_monitor()
assert first not in gc.callbacks
second = frame_timing.install_gc_monitor()
assert second is not first
assert sum(1 for cb in gc.callbacks if cb is second) == 1
def test_the_gc_monitor_needs_no_module_globals(monkeypatch):
# At shutdown, module globals can be torn down to None while a collection
# still calls the monitor ("'NoneType' object has no attribute
# 'perf_counter'"). Its clock is bound at construction.
monitor = frame_timing.GcMonitor(threshold=0.0)
monkeypatch.setattr(frame_timing, "time", None)
monkeypatch.setattr(frame_timing, "sys", None)
monitor("start", {"generation": 2})
monitor("stop", {"generation": 2})
assert monitor.collections == [0, 0, 1]
def test_the_gc_monitor_does_nothing_once_the_interpreter_is_finalizing(monkeypatch):
monitor = frame_timing.GcMonitor(threshold=0.0)
monkeypatch.setattr(monitor, "_is_finalizing", lambda: True)
monkeypatch.setattr(monitor, "_clock", lambda: pytest.fail("clock read"))
monitor("start", {"generation": 2})
monitor("stop", {"generation": 2})
assert monitor.collections == [0, 0, 0]
_EXIT_SCRIPT = """
import atexit, gc, sys
sys.path.insert(0, {root!r})
from src.common import frame_timing
# Registered before the monitor, so it runs after the monitor's own exit hook.
atexit.register(lambda: print("installed at exit:", any(
isinstance(cb, frame_timing.GcMonitor) for cb in gc.callbacks)))
monitor = frame_timing.install_gc_monitor()
gc.collect()
assert monitor.collections[2] >= 1
class Garbage:
# Makes cyclic garbage while modules are being torn down, so collections
# run during finalization.
def __del__(self):
for _ in range(5000):
cycle = []
cycle.append(cycle)
keep = Garbage()
keep.self = keep
"""
def test_a_process_with_the_gc_monitor_exits_cleanly():
import subprocess
root = str(Path(__file__).resolve().parent.parent)
proc = subprocess.run(
[sys.executable, "-c", _EXIT_SCRIPT.format(root=root)],
capture_output=True, text=True, timeout=60)
assert proc.returncode == 0, proc.stderr
assert "Exception ignored" not in proc.stderr
assert "installed at exit: False" in proc.stdout
+343
View File
@@ -0,0 +1,343 @@
"""The installer accepts Raspberry Pi OS Bookworm and Trixie, and nothing else.
Bookworm (Debian 12) ships Python 3.11 and Trixie (Debian 13) Python 3.13.
The rules live in scripts/install/lib_os.sh, which first_time_install.sh and
scripts/check_system_compatibility.sh both source. These tests feed the real
scripts a fake /etc/os-release (LM_OS_RELEASE_FILE) and stub python3,
systemctl and dpkg, so they need a Linux bash; the installer itself cannot
run end to end off a Pi.
"""
import importlib.util
import re
import shutil
import subprocess
import sys
from pathlib import Path
import pytest
ROOT = Path(__file__).resolve().parent.parent
LIB = ROOT / "scripts" / "install" / "lib_os.sh"
FIRST_TIME = ROOT / "first_time_install.sh"
COMPAT = ROOT / "scripts" / "check_system_compatibility.sh"
needs_bash = pytest.mark.skipif(
sys.platform == "win32" or shutil.which("bash") is None,
reason="runs the installer's shell code; needs a Linux bash",
)
OS_RELEASES = {
# Raspberry Pi OS 64-bit reports ID=debian, 32-bit ID=raspbian.
"trixie": 'PRETTY_NAME="Debian GNU/Linux 13 (trixie)"\nNAME="Debian GNU/Linux"\n'
'VERSION_ID="13"\nVERSION="13 (trixie)"\nVERSION_CODENAME=trixie\nID=debian\n',
"bookworm": 'PRETTY_NAME="Raspbian GNU/Linux 12 (bookworm)"\nNAME="Raspbian GNU/Linux"\n'
'VERSION_ID="12"\nVERSION="12 (bookworm)"\nVERSION_CODENAME=bookworm\nID=raspbian\n'
'ID_LIKE=debian\n',
"bookworm64": 'PRETTY_NAME="Debian GNU/Linux 12 (bookworm)"\nVERSION_ID="12"\n'
"VERSION_CODENAME=bookworm\nID=debian\n",
"bullseye": 'PRETTY_NAME="Raspbian GNU/Linux 11 (bullseye)"\nVERSION_ID="11"\n'
"VERSION_CODENAME=bullseye\nID=raspbian\n",
"ubuntu": 'PRETTY_NAME="Ubuntu 24.04 LTS"\nVERSION_ID="24.04"\nVERSION_CODENAME=noble\n'
"ID=ubuntu\nID_LIKE=debian\n",
"no-version-id": "PRETTY_NAME='Debian GNU/Linux trixie'\nVERSION_CODENAME=trixie\nID=debian\n",
}
def _stub(bin_dir: Path, name: str, body: str) -> None:
path = bin_dir / name
path.write_text("#!/bin/sh\n" + body, encoding="utf-8", newline="\n")
path.chmod(0o755)
def _stubs(tmp_path: Path, python_version="3.11", network="NetworkManager") -> Path:
"""python3 reports ``python_version`` (None: not installed); systemctl
reports ``network`` as the only active unit; dpkg lists no desktop."""
bin_dir = tmp_path / "bin"
bin_dir.mkdir(exist_ok=True)
if python_version is None:
# Shadows any real python3 further down PATH.
_stub(bin_dir, "python3", "exit 127\n")
else:
_stub(bin_dir, "python3", f'case "$*" in *"%d.%d.%d"*) echo "{python_version}.1" ;; '
f'*) echo "{python_version}" ;; esac\n')
_stub(bin_dir, "systemctl",
f'case "$*" in *"is-active --quiet {network}") exit 0 ;; esac\nexit 3\n')
_stub(bin_dir, "dpkg", "exit 0\n")
_stub(bin_dir, "dpkg-query", "exit 1\n")
_stub(bin_dir, "ping", "exit 0\n")
return bin_dir
def _env(tmp_path: Path, release: str, bin_dir: Path) -> dict:
os_release = tmp_path / "os-release"
os_release.write_text(OS_RELEASES[release], encoding="utf-8", newline="\n")
return {
"PATH": f"{bin_dir}:/usr/bin:/bin:/usr/sbin:/sbin",
"LM_OS_RELEASE_FILE": str(os_release),
"HOME": str(tmp_path),
}
def lib(snippet: str, env: dict) -> subprocess.CompletedProcess:
return subprocess.run(["bash", "-c", f"set -Eeuo pipefail\n. '{LIB}'\n{snippet}"],
capture_output=True, text=True, env=env)
# --- lib_os.sh -----------------------------------------------------------------
@needs_bash
class TestLibrary:
def test_library_is_syntactically_valid(self):
result = subprocess.run(["bash", "-n", str(LIB)], capture_output=True, text=True)
assert result.returncode == 0, result.stderr
@pytest.mark.parametrize("release,expected", [
("trixie", "trixie"),
("bookworm", "bookworm"),
("bookworm64", "bookworm"),
("no-version-id", "trixie"),
])
def test_supported_releases_are_recognised(self, tmp_path, release, expected):
result = lib("lm_os_release", _env(tmp_path, release, _stubs(tmp_path)))
assert result.returncode == 0, result.stderr
assert result.stdout.strip() == expected
@pytest.mark.parametrize("release", ["bullseye", "ubuntu"])
def test_other_systems_are_refused(self, tmp_path, release):
result = lib("lm_os_release", _env(tmp_path, release, _stubs(tmp_path)))
assert result.returncode != 0
assert result.stdout.strip() == ""
def test_missing_os_release_is_refused_not_fatal_to_the_caller(self, tmp_path):
env = _env(tmp_path, "trixie", _stubs(tmp_path))
env["LM_OS_RELEASE_FILE"] = str(tmp_path / "absent")
result = lib('if lm_os_release; then echo yes; else echo no; fi', env)
assert result.returncode == 0, result.stderr
assert result.stdout.strip() == "no"
def test_quoted_fields_are_unquoted(self, tmp_path):
env = _env(tmp_path, "no-version-id", _stubs(tmp_path))
result = lib("lm_os_field PRETTY_NAME; lm_os_field VERSION_ID", env)
assert result.stdout == "Debian GNU/Linux trixie\n"
@pytest.mark.parametrize("release,python", [("bookworm", "3.11"), ("trixie", "3.13")])
def test_each_release_names_the_python_it_ships(self, tmp_path, release, python):
result = lib(f"lm_release_python {release}", _env(tmp_path, "trixie", _stubs(tmp_path)))
assert result.stdout.strip() == python
@pytest.mark.parametrize("version,verdict", [
("3.11", "ok"), ("3.12", "ok"), ("3.13", "ok"),
("3.10", "too-old"), ("3.9", "too-old"), ("2.7", "too-old"),
("3.14", "too-new"), ("4.0", "too-new"),
("", "unknown"), ("garbage", "unknown"),
])
def test_python_versions(self, tmp_path, version, verdict):
result = lib(f'lm_python_check "{version}"', _env(tmp_path, "trixie", _stubs(tmp_path)))
assert result.returncode == 0, result.stderr
assert result.stdout.strip() == verdict
@pytest.mark.parametrize("active,expected", [
("NetworkManager", "networkmanager"),
("dhcpcd", "dhcpcd"),
("nothing", "unknown"),
])
def test_network_stack(self, tmp_path, active, expected):
env = _env(tmp_path, "bookworm", _stubs(tmp_path, network=active))
assert lib("lm_network_stack", env).stdout.strip() == expected
# --- first_time_install.sh's OS check ------------------------------------------
def _os_check_section() -> str:
"""first_time_install.sh from the OS check up to the next section, with
the desktop-marker directories pointed somewhere that cannot exist."""
text = FIRST_TIME.read_text(encoding="utf-8").replace("\r\n", "\n")
start = text.index("# Check OS version")
end = text.index("# The user who ran the installer")
section = text[start:end]
for marker in ("/usr/share/raspberrypi-ui-mods", "/usr/share/xsessions"):
assert marker in section
section = section.replace(marker, "/nonexistent" + marker)
return section
def run_os_check(tmp_path: Path, release: str, **stub_args) -> subprocess.CompletedProcess:
"""Run the OS check as the installer would, from a copy of the project
layout so ``$(dirname "$0")/scripts/install/lib_os.sh`` resolves."""
project = tmp_path / "project"
(project / "scripts" / "install").mkdir(parents=True)
shutil.copy(LIB, project / "scripts" / "install" / "lib_os.sh")
script = project / "first_time_install.sh"
script.write_text("set -Eeuo pipefail\n"
"trap 'echo ERR-TRAP line $LINENO >&2; exit 99' ERR\n"
+ _os_check_section() + '\necho "SECTION-DONE"\n',
encoding="utf-8", newline="\n")
env = _env(tmp_path, release, _stubs(tmp_path, **stub_args))
return subprocess.run(["bash", str(script)], capture_output=True, text=True, env=env)
@needs_bash
class TestInstallerOsCheck:
@pytest.mark.parametrize("release,python,label", [
("bookworm", "3.11", "Debian 12 (Bookworm)"),
("bookworm64", "3.11", "Debian 12 (Bookworm)"),
("trixie", "3.13", "Debian 13 (Trixie)"),
])
def test_supported_release_passes(self, tmp_path, release, python, label):
result = run_os_check(tmp_path, release, python_version=python)
assert result.returncode == 0, result.stdout + result.stderr
assert f"✓ {label} detected" in result.stdout
assert f"✓ Python {python} detected" in result.stdout
assert "✓ OS requirements met" in result.stdout
assert "SECTION-DONE" in result.stdout
@pytest.mark.parametrize("release,reason", [
("bullseye", "This version of Raspberry Pi OS is not supported"),
("ubuntu", "This script requires Raspberry Pi OS"),
])
def test_unsupported_system_stops_with_directions(self, tmp_path, release, reason):
result = run_os_check(tmp_path, release)
assert result.returncode == 1, result.stdout + result.stderr
assert reason in result.stdout
assert "Installation cannot continue." in result.stdout
assert "Trixie (Debian 13) or Bookworm (Debian 12)" in result.stdout
assert "SECTION-DONE" not in result.stdout
def test_python_older_than_the_rgbmatrix_floor_stops(self, tmp_path):
result = run_os_check(tmp_path, "bookworm", python_version="3.10")
assert result.returncode == 1, result.stdout + result.stderr
assert "needs Python 3.11 or newer" in result.stdout
assert "ships Python 3.11" in result.stdout
def test_untested_newer_python_warns_and_continues(self, tmp_path):
result = run_os_check(tmp_path, "trixie", python_version="3.14")
assert result.returncode == 0, result.stdout + result.stderr
assert "has not been tested with" in result.stdout
def test_missing_python_is_left_to_step_1(self, tmp_path):
result = run_os_check(tmp_path, "trixie", python_version=None)
assert result.returncode == 0, result.stdout + result.stderr
assert "Step 1 installs it" in result.stdout
def test_dhcpcd_is_explained_but_not_fatal(self, tmp_path):
result = run_os_check(tmp_path, "bookworm", network="dhcpcd")
assert result.returncode == 0, result.stdout + result.stderr
assert "dhcpcd, not NetworkManager" in result.stdout
assert "Network Config" in result.stdout
def test_networkmanager_is_confirmed(self, tmp_path):
result = run_os_check(tmp_path, "trixie", python_version="3.13")
assert "✓ NetworkManager is managing the network" in result.stdout
# --- check_system_compatibility.sh ---------------------------------------------
def run_compat(tmp_path: Path, release: str, **stub_args) -> subprocess.CompletedProcess:
env = _env(tmp_path, release, _stubs(tmp_path, **stub_args))
return subprocess.run(["bash", str(COMPAT)], capture_output=True, text=True, env=env)
@needs_bash
class TestCompatibilityCheck:
@pytest.mark.parametrize("release,python,label", [
("bookworm", "3.11", "Debian 12 (Bookworm)"),
("trixie", "3.13", "Debian 13 (Trixie)"),
])
def test_supported_release_is_reported_supported(self, tmp_path, release, python, label):
out = run_compat(tmp_path, release, python_version=python).stdout
assert f"Detected {label} - supported" in out
assert "Python version is supported (3.11-3.13)" in out
assert "not supported" not in out
assert "ships Python" not in out
@pytest.mark.parametrize("release", ["bullseye", "ubuntu"])
def test_unsupported_release_is_an_error(self, tmp_path, release):
result = run_compat(tmp_path, release)
assert result.returncode == 1
assert "is not supported - the installer requires Raspberry Pi OS Lite" in result.stdout
def test_python_below_the_floor_is_an_error(self, tmp_path):
result = run_compat(tmp_path, "bookworm", python_version="3.10")
assert result.returncode == 1
assert "Python 3.10 is too old - Python 3.11+ is required" in result.stdout
def test_dhcpcd_is_a_warning(self, tmp_path):
out = run_compat(tmp_path, "bookworm", network="dhcpcd").stdout
assert "dhcpcd manages the network" in out
assert "NetworkManager manages the network" not in out
def test_python_that_is_not_the_releases_own_is_flagged(self, tmp_path):
out = run_compat(tmp_path, "trixie", python_version="3.11").stdout
assert "Debian 13 (Trixie) ships Python 3.13, but python3 runs 3.11" in out
# --- things that must hold for both Python versions ----------------------------
def test_installer_scripts_do_not_hard_code_a_python_minor_version():
"""The services run /usr/bin/python3, which is 3.11 on Bookworm and 3.13
on Trixie; naming either one in an installer or a unit breaks the other."""
paths = [FIRST_TIME, *sorted((ROOT / "scripts" / "install").glob("*.sh")),
*sorted((ROOT / "systemd").glob("*.service"))]
offenders = []
for path in paths:
for n, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
if line.lstrip().startswith("#"):
continue
if re.search(r"python3\.1[0-9]", line):
offenders.append(f"{path.relative_to(ROOT)}:{n}: {line.strip()}")
assert not offenders, "\n".join(offenders)
def _load_apt_installer():
spec = importlib.util.spec_from_file_location(
"install_dependencies_apt", ROOT / "scripts" / "install_dependencies_apt.py")
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
class TestAptFallbackRespectsThePins:
"""Step 7's apt-first fallback must not accept the releases' older apt
copies (Bookworm: Flask 2.2.2, Pillow 9.4; Trixie: Flask 3.1.1)."""
def test_floors_come_from_the_web_requirements(self):
mod = _load_apt_installer()
for name in ("flask", "werkzeug", "pillow", "requests", "psutil", "pytz", "freetype-py"):
assert mod.MIN_VERSIONS.get(name), f"no floor read for {name}"
assert mod.MIN_VERSIONS["freetype-py"] == (2, 5, 1)
@pytest.mark.parametrize("package,apt_version", [
("flask", (2, 2, 2)), # Bookworm
("flask", (3, 1, 1)), # Trixie
("PIL", (9, 4, 0)), # Bookworm python3-pil
("werkzeug", (2, 2, 2)),
("freetype-py", (2, 3, 0)),
])
def test_an_apt_copy_below_the_pin_does_not_count(self, monkeypatch, package, apt_version):
mod = _load_apt_installer()
monkeypatch.setattr(mod, "_installed_version_tuple", lambda dist: apt_version)
monkeypatch.setitem(sys.modules, mod.IMPORT_NAME_MAP.get(package, package), object())
assert mod.check_package_installed(package) is False
def test_a_version_at_the_pin_counts(self, monkeypatch):
mod = _load_apt_installer()
floor = mod.MIN_VERSIONS["flask"]
monkeypatch.setattr(mod, "_installed_version_tuple", lambda dist: floor)
monkeypatch.setitem(sys.modules, "flask", object())
assert mod.check_package_installed("flask") is True
def test_the_pillow_version_is_looked_up_under_its_distribution_name(self, monkeypatch):
mod = _load_apt_installer()
seen = []
monkeypatch.setattr(mod, "_installed_version_tuple", lambda dist: seen.append(dist) or (99,))
monkeypatch.setitem(sys.modules, "PIL", object())
assert mod.check_package_installed("PIL") is True
assert seen == ["Pillow"]
def test_pip_is_asked_for_pillow_not_pil(self, monkeypatch):
mod = _load_apt_installer()
calls = []
monkeypatch.setattr(mod, "_run", lambda cmd: calls.append(cmd) or (True, ""))
assert mod.install_via_pip("PIL") == (True, "")
assert calls[0][-1] == "Pillow"
+2 -2
View File
@@ -6,8 +6,8 @@ Complete / Web UI Access" summary. `reboot` returns at once and the script
carried on printing while the system went down, so the SSH session usually carried on printing while the system went down, so the SSH session usually
dropped before the user saw the web UI address. dropped before the user saw the web UI address.
first_time_install.sh exits on anything but Raspberry Pi OS Trixie before it first_time_install.sh exits on anything but Raspberry Pi OS Bookworm or Trixie
parses its arguments, so the behavioural test runs only the tail of the before it parses its arguments, so the behavioural test runs only the tail of the
script -- from the summary to the end -- with systemctl, nmcli, hostname, ip script -- from the summary to the end -- with systemctl, nmcli, hostname, ip
and reboot stubbed. and reboot stubbed.
""" """
+370
View File
@@ -0,0 +1,370 @@
"""New installs run the newest release; re-running the installer never moves backwards.
#684 made devices update along a channel -- stable follows the newest vX.Y.Z
tag, beta follows main -- but a new install still cloned main's tip, so it ran
unreleased code until the next release caught up with it. The one-shot
installer (scripts/install/one-shot-install.sh, which is where the clone
happens) now checks out the newest release after cloning, unless
LEDMATRIX_CHANNEL=beta. Re-running it on an existing checkout moves a stable
device forward to the newest release only when that release contains its
commit, as update_channel.checkout_release() does, and leaves beta devices
(and stable ones newer than every release) on the fast-forward pull they
always had. first_time_install.sh writes an explicitly chosen channel
(--beta / LEDMATRIX_CHANNEL) into config.json.
These run the installer's own bash, under its strict mode, against real git
repositories.
"""
import json
import re
import subprocess
import sys
from pathlib import Path
import pytest
ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(ROOT))
ONE_SHOT = ROOT / "scripts" / "install" / "one-shot-install.sh"
INSTALLER = ROOT / "first_time_install.sh"
BEGIN = "# --- release checkout helpers"
END = "# --- end release checkout helpers"
from web_interface import update_channel # noqa: E402
pytestmark = pytest.mark.skipif(
not sys.platform.startswith("linux"), reason="runs the installer's bash under Linux"
)
def helper_block() -> str:
text = ONE_SHOT.read_text(encoding="utf-8")
assert text.count(BEGIN) == 1 and text.count(END) == 1, "helper block markers missing or duplicated"
return text[text.index(BEGIN): text.index(END)]
def git(*args, cwd, env):
result = subprocess.run(["git", *args], cwd=cwd, env=env, capture_output=True, text=True)
assert result.returncode == 0, f"git {' '.join(args)} failed: {result.stderr}"
return result.stdout.strip()
@pytest.fixture
def git_env(tmp_path):
config = tmp_path / "gitconfig"
config.write_text(
"[user]\n\tname = t\n\temail = t@t\n"
"[protocol \"file\"]\n\tallow = always\n"
"[init]\n\tdefaultBranch = main\n"
"[advice]\n\tdetachedHead = false\n",
encoding="utf-8",
)
return {
"PATH": "/usr/bin:/bin:/usr/sbin:/sbin",
"HOME": str(tmp_path),
"GIT_CONFIG_GLOBAL": str(config),
"GIT_CONFIG_NOSYSTEM": "1",
}
#: Tags and the commit (index into the history) each points at. The newest
#: release is v3.10.0: 10 > 8 numerically, the rc and the zero-padded tag are
#: not releases, and v3.12 and nightly are not vX.Y.Z at all.
TAGS = {
"v3.7.0": 0, "v3.8.0": 1, "v3.10.0": 2,
"v3.11.0-rc1": 3, "v03.12.0": 3, "v3.12": 3, "nightly": 3,
}
NEWEST = "v3.10.0"
@pytest.fixture
def origin(tmp_path, git_env):
"""A stand-in for GitHub: five commits on main (the last newer than any release)."""
seed = tmp_path / "seed"
seed.mkdir()
git("init", "-q", ".", cwd=seed, env=git_env)
commits = []
for i in range(5):
(seed / "version.txt").write_text(str(i), encoding="utf-8")
git("add", ".", cwd=seed, env=git_env)
git("commit", "-qm", f"c{i}", cwd=seed, env=git_env)
commits.append(git("rev-parse", "HEAD", cwd=seed, env=git_env))
for tag, index in TAGS.items():
git("tag", tag, commits[index], cwd=seed, env=git_env)
bare = tmp_path / "origin.git"
git("clone", "-q", "--bare", str(seed), str(bare), cwd=tmp_path, env=git_env)
return bare, commits, seed
def run_block(snippet, cwd, env, channel=None):
env = dict(env)
if channel is not None:
env["LEDMATRIX_CHANNEL"] = channel
script = (
"set -Eeuo pipefail\n"
"trap 'echo ERR_TRAP_FIRED >&2; exit 99' ERR\n"
'print_success() { echo "OK: $*"; }\n'
'print_warning() { echo "W: $*"; }\n'
f"{helper_block()}\n"
f"{snippet}\n"
)
result = subprocess.run(["bash", "-c", script], cwd=cwd, capture_output=True, text=True, env=env)
assert "ERR_TRAP_FIRED" not in result.stderr, result.stdout + result.stderr
return result
def clone(origin, tmp_path, env, name="LEDMatrix"):
bare, _, _ = origin
target = tmp_path / name
git("clone", "-q", str(bare), str(target), cwd=tmp_path, env=env)
return target
def head(repo, env):
return git("rev-parse", "HEAD", cwd=repo, env=env)
def branch(repo, env):
result = subprocess.run(["git", "symbolic-ref", "--quiet", "--short", "HEAD"], cwd=repo, env=env,
capture_output=True, text=True)
return result.stdout.strip()
def set_channel(repo, channel):
(repo / "config").mkdir(exist_ok=True)
(repo / "config" / "config.json").write_text(
json.dumps({"auto_update": {"enabled": False, "channel": channel}}), encoding="utf-8")
# -- a fresh install -------------------------------------------------------------
def test_a_fresh_clone_checks_out_the_newest_release(origin, tmp_path, git_env):
_, commits, _ = origin
repo = clone(origin, tmp_path, git_env)
out = run_block("_lm_checkout_release_after_clone", repo, git_env)
assert out.returncode == 0
assert head(repo, git_env) == commits[TAGS[NEWEST]]
assert branch(repo, git_env) == "", "a release is checked out detached, as Update Code does"
assert f"Installing release {NEWEST}" in out.stdout
@pytest.mark.parametrize("channel", ["beta", "BETA", " beta "])
def test_a_fresh_beta_install_stays_on_main(origin, tmp_path, git_env, channel):
_, commits, _ = origin
repo = clone(origin, tmp_path, git_env)
run_block("_lm_checkout_release_after_clone", repo, git_env, channel=channel)
assert head(repo, git_env) == commits[-1] and branch(repo, git_env) == "main"
def test_an_unknown_channel_falls_back_to_stable_and_says_so(origin, tmp_path, git_env):
_, commits, _ = origin
repo = clone(origin, tmp_path, git_env)
out = run_block("_lm_checkout_release_after_clone", repo, git_env, channel="nightly")
assert head(repo, git_env) == commits[TAGS[NEWEST]]
assert "not stable or beta" in out.stdout + out.stderr
def test_a_repository_without_releases_installs_main(tmp_path, git_env):
seed = tmp_path / "seed"
seed.mkdir()
git("init", "-q", ".", cwd=seed, env=git_env)
(seed / "f").write_text("x", encoding="utf-8")
git("add", ".", cwd=seed, env=git_env)
git("commit", "-qm", "only", cwd=seed, env=git_env)
git("tag", "v3.0", cwd=seed, env=git_env) # not a release tag
repo = tmp_path / "LEDMatrix"
git("clone", "-q", str(seed), str(repo), cwd=tmp_path, env=git_env)
out = run_block("_lm_checkout_release_after_clone", repo, git_env)
assert branch(repo, git_env) == "main" and "No release found" in out.stdout
# -- re-running on an existing checkout -----------------------------------------------
def existing(origin, tmp_path, env, at, detached):
"""An installed checkout, at commit index ``at``, on main or detached."""
_, commits, _ = origin
repo = clone(origin, tmp_path, env)
if detached:
git("checkout", "-q", "--detach", commits[at], cwd=repo, env=env)
else:
git("reset", "-q", "--hard", commits[at], cwd=repo, env=env)
return repo
def update(repo, env, channel=None):
out = run_block("if _lm_update_existing_checkout; then echo RESULT=handled; "
"else echo RESULT=pull; fi", repo, env, channel=channel)
return re.search(r"RESULT=(\w+)", out.stdout).group(1), out
def test_a_device_on_an_older_release_moves_to_the_newest(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS["v3.7.0"], detached=True)
result, out = update(repo, git_env)
assert result == "handled"
assert head(repo, git_env) == commits[TAGS[NEWEST]]
assert f"Updated to release {NEWEST}" in out.stdout
def test_a_device_already_on_the_newest_release_stays(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS[NEWEST], detached=True)
result, out = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[TAGS[NEWEST]]
assert "Already on the newest release" in out.stdout
def test_a_device_on_main_behind_the_newest_release_moves_to_it(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS["v3.8.0"], detached=False)
result, _ = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[TAGS[NEWEST]]
def test_a_device_on_main_newer_than_every_release_is_not_moved_back(origin, tmp_path, git_env):
"""It keeps the fast-forward pull it always had, and waits for a release to contain it."""
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=4, detached=False)
result, _ = update(repo, git_env)
assert result == "pull"
assert head(repo, git_env) == commits[4] and branch(repo, git_env) == "main"
def test_a_detached_device_newer_than_every_release_is_left_alone(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=3, detached=True)
result, out = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[3]
assert "newer than the newest release" in out.stdout
def test_a_higher_version_on_an_older_commit_is_not_a_downgrade(origin, tmp_path, git_env):
"""Newest by version is not newest by history: never move to a tag that does not contain HEAD."""
bare, commits, seed = origin
git("tag", "v9.0.0", commits[0], cwd=seed, env=git_env)
git("push", "-q", str(bare), "v9.0.0", cwd=seed, env=git_env)
repo = existing(origin, tmp_path, git_env, at=TAGS[NEWEST], detached=True)
result, _ = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[TAGS[NEWEST]]
def test_a_new_release_published_since_the_clone_is_fetched(origin, tmp_path, git_env):
bare, commits, seed = origin
repo = existing(origin, tmp_path, git_env, at=TAGS[NEWEST], detached=True)
git("tag", "v3.11.0", commits[4], cwd=seed, env=git_env)
git("push", "-q", str(bare), "v3.11.0", cwd=seed, env=git_env)
result, _ = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[4]
def test_a_beta_device_keeps_its_pull(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS["v3.8.0"], detached=False)
set_channel(repo, "beta")
result, _ = update(repo, git_env)
assert result == "pull" and head(repo, git_env) == commits[TAGS["v3.8.0"]]
def test_the_environment_overrides_the_configured_channel(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS["v3.8.0"], detached=False)
set_channel(repo, "stable")
result, _ = update(repo, git_env, channel="beta")
assert result == "pull" and head(repo, git_env) == commits[TAGS["v3.8.0"]]
def test_local_edits_that_block_the_move_keep_the_checkout(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS["v3.7.0"], detached=True)
(repo / "version.txt").write_text("my edit", encoding="utf-8")
result, out = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[TAGS["v3.7.0"]]
assert (repo / "version.txt").read_text(encoding="utf-8") == "my edit"
assert "Could not move to release" in out.stdout
def test_an_unreachable_origin_keeps_the_checkout(origin, tmp_path, git_env):
_, commits, _ = origin
repo = existing(origin, tmp_path, git_env, at=TAGS["v3.7.0"], detached=True)
git("remote", "set-url", "origin", str(tmp_path / "gone.git"), cwd=repo, env=git_env)
result, out = update(repo, git_env)
assert result == "handled" and head(repo, git_env) == commits[TAGS["v3.7.0"]]
assert "Could not fetch" in out.stdout
# -- same rules as the web interface -------------------------------------------------
NAMES = ["v1.2.3", "v1.10.0", "v1.9.9", "v2.0.0-rc1", "v02.0.0", "v2.0", "v10.0.0", "v9.99.99",
"release-11", "v10.0.0+build", "v0.0.0", "v10.0.1", "v1.2.03", "V11.0.0"]
@pytest.mark.parametrize("subset", [NAMES, NAMES[:4], ["v2.0", "nightly"], NAMES[::-1][:6]])
def test_the_newest_tag_matches_update_channel(tmp_path, git_env, subset):
repo = tmp_path / "tags"
repo.mkdir()
git("init", "-q", ".", cwd=repo, env=git_env)
git("commit", "-q", "--allow-empty", "-m", "x", cwd=repo, env=git_env)
for name in subset:
git("tag", name, cwd=repo, env=git_env)
out = run_block("_lm_newest_release_tag", repo, git_env).stdout.strip()
assert out == (update_channel.newest_release_tag(subset) or "")
# -- wiring --------------------------------------------------------------------------
def test_every_clone_is_followed_by_the_release_checkout():
text = ONE_SHOT.read_text(encoding="utf-8")
clones = [m.start() for m in re.finditer(r'retry git clone "\$REPO_URL" "\$REPO_DIR"\n', text)]
assert clones
for pos in clones:
following = text[pos:].splitlines()[1]
assert '_lm_checkout_release_after_clone' in following, following
def test_the_existing_checkout_is_handled_before_the_old_pull():
text = ONE_SHOT.read_text(encoding="utf-8")
assert re.search(r'if _lm_update_existing_checkout; then\n\s+PULL_SUCCESS=true\n'
r'\s+elif git pull --ff-only origin "\$CURRENT_BRANCH"', text)
def test_the_one_shot_passes_the_channel_to_the_installer():
text = ONE_SHOT.read_text(encoding="utf-8")
assert 'LEDMATRIX_CHANNEL="${LEDMATRIX_CHANNEL:-}"' in text
# -- first_time_install.sh records the chosen channel ----------------------------------
CHANNEL_BEGIN = 'case "$UPDATE_CHANNEL" in'
CHANNEL_END = 'set it from the General tab instead"\n fi\nfi\n'
def channel_block():
text = INSTALLER.read_text(encoding="utf-8")
start = text.index(CHANNEL_BEGIN)
return text[start: text.index(CHANNEL_END, start) + len(CHANNEL_END)]
@pytest.mark.parametrize("auto_update, channel, expected", [
("", "beta", {"enabled": False, "channel": "beta"}),
("", "stable", {"enabled": False, "channel": "stable"}),
("1", "", {"enabled": True, "channel": "stable"}),
("", "", {"enabled": False, "channel": "stable"}), # nothing asked: untouched
("", "nightly", {"enabled": False, "channel": "stable"}), # nonsense: untouched
])
def test_the_installer_writes_only_an_explicit_channel(tmp_path, auto_update, channel, expected):
(tmp_path / "config").mkdir()
config = tmp_path / "config" / "config.json"
config.write_text(json.dumps({"auto_update": {"enabled": False, "channel": "stable"}, "x": 1}))
script = (f'set -Eeuo pipefail\nPROJECT_ROOT_DIR="{tmp_path}"\nAUTO_UPDATE="{auto_update}"\n'
f'UPDATE_CHANNEL="{channel}"\n{channel_block()}')
result = subprocess.run(["bash", "-c", script], capture_output=True, text=True)
assert result.returncode == 0, result.stdout + result.stderr
data = json.loads(config.read_text())
assert data["auto_update"] == expected and data["x"] == 1
def test_the_installer_accepts_beta_as_a_flag_and_from_the_environment():
text = INSTALLER.read_text(encoding="utf-8")
assert re.search(r"^\s*--beta\) UPDATE_CHANNEL=beta ;;", text, re.M)
assert 'UPDATE_CHANNEL=$(printf \'%s\' "${LEDMATRIX_CHANNEL:-}"' in text
assert "LEDMATRIX_CHANNEL=stable|beta" in text, "documented in --help"
+96
View File
@@ -236,3 +236,99 @@ class TestOneShotContract:
assert "--recurse-submodules" not in ONE_SHOT.read_text(encoding="utf-8") assert "--recurse-submodules" not in ONE_SHOT.read_text(encoding="utf-8")
installer = INSTALLER.read_text(encoding="utf-8") installer = INSTALLER.read_text(encoding="utf-8")
assert re.search(rf"submodule update --init --recursive {SUB}", installer) assert re.search(rf"submodule update --init --recursive {SUB}", installer)
PATCH_DIR = ROOT / "patches" / "rpi-rgb-led-matrix"
def _write_patch(project: Path, name: str, old: str, new: str) -> Path:
"""A git patch rewriting the fake library's Makefile from `old` to `new`."""
patch = project / "patches" / "rpi-rgb-led-matrix" / name
patch.parent.mkdir(parents=True, exist_ok=True)
patch.write_text(
"A note before the diff, as the shipped patches carry.\n\n"
"diff --git a/Makefile b/Makefile\n"
"--- a/Makefile\n"
"+++ b/Makefile\n"
"@@ -1 +1 @@\n"
f"-{old}\n"
"\\ No newline at end of file\n"
f"+{new}\n"
"\\ No newline at end of file\n",
encoding="utf-8",
)
return patch
def _makefile(project: Path) -> str:
return (project / SUB / "Makefile").read_text(encoding="utf-8")
def _status(project: Path, env: dict) -> str:
return git("status", "--porcelain", cwd=project / SUB, env=env)
class TestLibraryPatches:
"""patches/rpi-rgb-led-matrix/*.patch go in for the build and come back out."""
def test_applied_for_the_build_and_reverted_after(self, make_project, git_env):
project = make_project("B")
_write_patch(project, "0001-x.patch", "A", "patched")
result = run_helpers(
f'_apply_rgb_patches; cat "{project / SUB / "Makefile"}"; echo; _revert_rgb_patches',
project, git_env)
assert result.returncode == 0, result.stderr
assert "Applied library patch 0001-x.patch" in result.stdout
assert "patched" in result.stdout # what the build would see
assert _makefile(project) == "A" # and the checkout afterwards
assert _status(project, git_env) == ""
def test_already_applied_is_left_alone(self, make_project, git_env):
project = make_project("B")
_write_patch(project, "0001-x.patch", "A", "patched")
(project / SUB / "Makefile").write_text("patched", encoding="utf-8")
result = run_helpers("_apply_rgb_patches; _revert_rgb_patches", project, git_env)
assert result.returncode == 0, result.stderr
assert "already applied" in result.stdout
assert _makefile(project) == "patched" # not reverted: it was not ours
def test_a_patch_that_does_not_apply_is_skipped_not_fatal(self, make_project, git_env):
project = make_project("B")
_write_patch(project, "0001-x.patch", "something else", "patched")
result = run_helpers("_apply_rgb_patches; _revert_rgb_patches; echo REACHED", project, git_env)
assert result.returncode == 0, result.stderr
assert "ERR_TRAP_FIRED" not in result.stderr
assert "does not apply" in result.stdout and "REACHED" in result.stdout
assert _makefile(project) == "A"
def test_no_patch_directory_is_a_no_op(self, make_project, git_env):
project = make_project("B")
result = run_helpers("_apply_rgb_patches; _revert_rgb_patches; echo REACHED", project, git_env)
assert result.returncode == 0, result.stderr
assert result.stdout.strip() == "REACHED"
def test_revert_twice_is_harmless(self, make_project, git_env):
# The EXIT trap runs _revert_rgb_patches again after Step 6 already has.
project = make_project("B")
_write_patch(project, "0001-x.patch", "A", "patched")
result = run_helpers("_apply_rgb_patches; _revert_rgb_patches; _revert_rgb_patches", project, git_env)
assert result.returncode == 0, result.stderr
assert _makefile(project) == "A" and "Could not revert" not in result.stdout
def test_build_is_wrapped_and_the_exit_trap_reverts(self):
installer = INSTALLER.read_text(encoding="utf-8")
apply_at = installer.index("_apply_rgb_patches\n if run_rgbmatrix_build")
assert installer.index("_revert_rgb_patches\n cat \"$BUILD_OUTPUT\"") > apply_at
assert re.search(r"^trap '[^']*_revert_rgb_patches[^']*' EXIT", installer, re.M)
def test_shipped_patches_are_git_patches_against_the_library(self, tmp_path, git_env):
patches = sorted(PATCH_DIR.glob("*.patch"))
assert patches, "no shipped library patches"
repo = tmp_path / "r"
repo.mkdir()
git("init", "-q", ".", cwd=repo, env=git_env)
for patch in patches:
stat = git("apply", "--numstat", str(patch), cwd=repo, env=git_env)
files = {line.split("\t")[2] for line in stat.splitlines()}
assert files <= {"lib/framebuffer.cc", "bindings/python/rgbmatrix/core.pyx",
"bindings/python/rgbmatrix/cppinc.pxd"}, files
+619
View File
@@ -0,0 +1,619 @@
"""Control socket stage 3: the display's state over the socket (state.get,
state.subscribe) instead of polled cache keys.
* ``StateHub`` and the request handling are plain Python, tested on every
platform: versions move only when something a reader sees changes, the
``since``/``epoch`` short answer, an oversized snapshot drops only its
plugin section, and publishing never waits for a reader.
* ``TestLiveStream`` needs AF_UNIX (Linux, WSL, a Pi; skipped on Windows): a
real server pushing to real subscribers -- several at once, one that never
reads, one whose display goes away and comes back.
"""
import json
import socket
import threading
import time
import pytest
from src.ipc import client
from src.ipc import contract as c
from src.ipc import server as srv
from src.ipc.contract import Command, ErrorCode
from src.ipc.server import ControlServer, StateHub, fit_snapshot
needs_unix_sockets = pytest.mark.skipif(not c.socket_supported(),
reason='AF_UNIX sockets are Linux/macOS only')
class FakeClock:
def __init__(self, now=1000.0):
self.now = now
def __call__(self):
return self.now
def _req(cmd, args=None, rid='r1', v=1):
return json.dumps({'v': v, 'id': rid, 'cmd': cmd, 'args': args or {}}).encode()
def _loop(age=1.0):
return lambda: {'heartbeat_age_seconds': age, 'armed': age is not None,
'stale_after': 60.0}
def _display(mode='clock', active=True, updated=None):
return {'mode': mode, 'plugin_id': mode, 'mode_index': 0, 'total_modes': 2,
'on_demand_active': False, 'is_display_active': active,
'last_updated': time.time() if updated is None else updated}
@pytest.fixture
def hub():
h = StateHub(loop_probe=_loop(), epoch='e1', pid=4242)
h.publish('display', _display(), volatile=('last_updated',))
return h
# --- the hub ---------------------------------------------------------------------
class TestStateHub:
def test_snapshot_has_every_section_and_the_envelope(self, hub):
snap = hub.snapshot()
assert snap['schema'] == c.STATE_SCHEMA
assert (snap['version'], snap['epoch'], snap['pid']) == (1, 'e1', 4242)
assert snap['changed'] is True
assert list(snap['state']) == list(c.STATE_SECTIONS)
assert snap['state']['display']['mode'] == 'clock'
assert snap['state']['on_demand'] is None # not published yet
assert snap['state']['loop'] == snap['loop'] == _loop()()
def test_only_a_real_change_is_a_new_version(self, hub):
assert not hub.publish('display', _display(updated=1.0), volatile=('last_updated',))
assert hub.version == 1
# ...but the latest timestamp is what a reader gets.
assert hub.snapshot()['state']['display']['last_updated'] == 1.0
assert hub.publish('display', _display(mode='weather'), volatile=('last_updated',))
assert hub.version == 2
assert hub.publish('brightness', {'brightness': 50})
assert not hub.publish('brightness', {'brightness': 50})
assert hub.version == 3
def test_the_published_dict_is_copied(self, hub):
value = {'brightness': 50}
hub.publish('brightness', value)
value['brightness'] = 10
assert hub.snapshot()['state']['brightness'] == {'brightness': 50}
def test_since_the_current_version_is_the_short_answer(self, hub):
short = hub.snapshot(since=1, epoch='e1')
assert short['changed'] is False and 'state' not in short
assert short['version'] == 1 and short['loop']['heartbeat_age_seconds'] == 1.0
# An older version, or another epoch (a restarted display): the full state.
assert hub.snapshot(since=0, epoch='e1')['changed'] is True
assert hub.snapshot(since=1, epoch='other')['changed'] is True
assert hub.snapshot(since=1)['changed'] is True
def test_the_short_answer_carries_the_latest_volatile_values(self, hub):
"""A version that stays put must not freeze the timestamps a reader
judges freshness by: the short answer (a tick) carries them."""
hub.publish('on_demand', {'active': False, 'last_updated': 2.0, 'remaining': 0.0},
volatile=('last_updated', 'remaining'))
hub.publish('brightness', {'brightness': 50})
version = hub.version
assert not hub.publish('display', _display(updated=1234.5), volatile=('last_updated',))
short = hub.snapshot(since=version, epoch='e1')
assert short['changed'] is False and 'state' not in short
assert short['volatile'] == {'display': {'last_updated': 1234.5},
'on_demand': {'last_updated': 2.0, 'remaining': 0.0}}
# The full answer has them in place already.
assert 'volatile' not in hub.snapshot()
def test_volatile_values_skip_sections_without_any(self):
hub = StateHub(epoch='e1')
hub.publish('brightness', {'brightness': 50})
hub.publish('plugins', None, volatile=('published_at',))
assert hub.snapshot(since=hub.version, epoch='e1')['volatile'] == {}
def test_each_display_run_has_its_own_epoch(self):
assert StateHub().epoch != StateHub().epoch
def test_the_loop_is_measured_when_asked(self):
ages = iter([1.0, 61.0])
hub = StateHub(loop_probe=lambda: {'heartbeat_age_seconds': next(ages),
'armed': True, 'stale_after': 60.0})
assert hub.snapshot()['loop']['heartbeat_age_seconds'] == 1.0
# Nothing was published -- a stuck render thread publishes nothing --
# and the age still moves.
assert hub.snapshot()['loop']['heartbeat_age_seconds'] == 61.0
def test_a_failing_probe_is_unknown_not_raised(self):
def boom():
raise RuntimeError('x')
assert StateHub(loop_probe=boom).loop()['heartbeat_age_seconds'] is None
assert StateHub().loop()['armed'] is False
def test_wait_for_change(self, hub):
assert hub.wait_for_change(1, 0.01) is False
threading.Timer(0.05, lambda: hub.publish('brightness', {'b': 1})).start()
start = time.monotonic()
assert hub.wait_for_change(1, 5.0) is True
assert time.monotonic() - start < 2.0
def test_wait_ends_on_stop(self, hub):
stop = threading.Event()
def end():
stop.set()
hub.wake()
threading.Timer(0.05, end).start()
start = time.monotonic()
assert hub.wait_for_change(1, 5.0, stop) is False
assert time.monotonic() - start < 2.0
def test_readers_active(self):
clock = FakeClock()
hub = StateHub(clock=clock, reader_window=60.0)
assert not hub.readers_active()
hub.note_read()
clock.now += 59
assert hub.readers_active()
clock.now += 2
assert not hub.readers_active()
hub.subscriber_joined()
clock.now += 1000
assert hub.readers_active() # for as long as it is connected
hub.subscriber_left()
assert hub.readers_active() # and one window after
clock.now += 61
assert not hub.readers_active()
def test_publishing_never_waits_for_readers(self, hub):
"""Many readers snapshotting and waiting at once: every publish from
the 'render thread' still returns in well under a frame."""
stop = threading.Event()
def reader():
version = 0
while not stop.is_set():
hub.wait_for_change(version, 0.01)
version = hub.snapshot()['version']
threads = [threading.Thread(target=reader, daemon=True) for _ in range(8)]
for t in threads:
t.start()
worst = 0.0
try:
for i in range(500):
start = time.perf_counter()
hub.publish('display', _display(mode=f'm{i}'), volatile=('last_updated',))
worst = max(worst, time.perf_counter() - start)
finally:
stop.set()
for t in threads:
t.join(2)
assert worst < 0.05
class TestFitSnapshot:
def test_small_is_unchanged(self, hub):
snap = hub.snapshot()
assert fit_snapshot(snap) is snap
def test_too_large_drops_only_the_plugins(self, hub):
plugins = {f'p{i}': {'error': {'message': 'x' * 200}} for i in range(400)}
hub.publish('plugins', {'schema': 1, 'plugins': plugins})
fitted = fit_snapshot(hub.snapshot())
assert fitted['state']['plugins'] is None
assert fitted['truncated'] == ['plugins']
assert fitted['state']['display']['mode'] == 'clock'
c.encode_message({'v': 1, 'id': 'x' * 128, 'ok': True, 'result': fitted})
# --- the commands, without a socket ------------------------------------------------
class TestHandleLine:
def test_state_get(self, hub):
server = ControlServer('/unused.sock', state_hub=hub)
resp = server.handle_line(_req(Command.STATE_GET))
assert resp.ok and resp.result['version'] == 1
assert resp.result['state']['display']['mode'] == 'clock'
assert hub.readers_active()
def test_state_get_since(self, hub):
server = ControlServer('/unused.sock', state_hub=hub)
resp = server.handle_line(_req(Command.STATE_GET, {'since': 1, 'epoch': 'e1'}))
assert resp.ok and resp.result['changed'] is False and 'state' not in resp.result
@pytest.mark.parametrize('args', [{'since': -1}, {'since': 'x'}, {'since': True},
{'epoch': 5}, {'epoch': 'x' * 200}])
def test_bad_args(self, hub, args):
server = ControlServer('/unused.sock', state_hub=hub)
resp = server.handle_line(_req(Command.STATE_GET, args))
assert not resp.ok and resp.error.code == ErrorCode.INVALID_ARGS
def test_without_a_hub_it_is_an_error_not_a_crash(self):
server = ControlServer('/unused.sock')
for cmd in (Command.STATE_GET, Command.STATE_SUBSCRIBE):
resp = server.handle_line(_req(cmd))
assert not resp.ok and resp.error.code == ErrorCode.INTERNAL
def test_hello_lists_the_new_commands(self, hub):
resp = ControlServer('/unused.sock', state_hub=hub).handle_line(
_req(Command.HELLO, {'versions': [1]}))
assert {Command.STATE_GET, Command.STATE_SUBSCRIBE} <= set(resp.result['commands'])
def test_state_commands_are_not_queued(self, hub):
server = ControlServer('/unused.sock', state_hub=hub)
server.handle_line(_req(Command.STATE_GET))
assert not server.has_pending and server.drain() == []
class TestEvents:
def test_round_trip(self):
event = c.StateEvent('s1', c.StateEventKind.TICK, {'version': 3})
obj = json.loads(c.encode_message(event.to_dict()))
assert c.is_event(obj)
assert c.StateEvent.from_dict(obj) == event
def test_a_response_is_not_an_event(self):
assert not c.is_event(c.Response.success('x', {}).to_dict())
@pytest.mark.parametrize('obj', [
[], {'v': 1, 'id': 's', 'event': 'other', 'result': {}},
{'v': 1, 'id': 's', 'event': 'state', 'result': []},
{'v': '1', 'id': 's', 'event': 'state', 'result': {}},
{'v': 1, 'id': None, 'event': 'state', 'result': {}},
])
def test_bad_events_are_refused(self, obj):
with pytest.raises(c.ProtocolError):
c.StateEvent.from_dict(obj)
class TestSubscriptionStore:
"""StateSubscription's bookkeeping, without a socket."""
def test_latest_is_none_until_a_snapshot_and_after_silence(self, hub):
clock = FakeClock()
sub = client.StateSubscription(paths=['/nowhere'], silence=15.0, clock=clock)
assert sub.latest() is None
sub._store(hub.snapshot(), full=True)
latest = sub.latest()
assert latest['version'] == 1 and latest['received_mono'] == clock.now
clock.now += 16
assert sub.latest() is None
def test_a_tick_refreshes_the_loop_and_keeps_the_state(self, hub):
clock = FakeClock()
sub = client.StateSubscription(paths=['/nowhere'], clock=clock)
sub._store(hub.snapshot(), full=True)
clock.now += 10
tick = {'version': 1, 'epoch': 'e1', 'served_at': 5.0,
'loop': {'heartbeat_age_seconds': 70.0, 'armed': True, 'stale_after': 60.0}}
sub._store(tick, full=False)
latest = sub.latest()
assert latest['state']['display']['mode'] == 'clock'
assert latest['state']['loop']['heartbeat_age_seconds'] == 70.0
assert latest['received_mono'] == clock.now
def test_a_tick_refreshes_the_volatile_timestamps(self, hub):
"""The bug from the ledpi rig: the same mode on screen for minutes
left the reader's display.last_updated at the last real change."""
clock = FakeClock()
sub = client.StateSubscription(paths=['/nowhere'], clock=clock)
sub._store(hub.snapshot(), full=True)
hub.publish('display', _display(updated=9999.0), volatile=('last_updated',))
sub._store(hub.snapshot(since=1, epoch='e1'), full=False)
latest = sub.latest()
assert latest['state']['display']['last_updated'] == 9999.0
assert latest['state']['display']['mode'] == 'clock'
assert latest['version'] == 1
def test_a_tick_adds_no_section_or_key_and_ignores_another_version(self, hub):
clock = FakeClock()
sub = client.StateSubscription(paths=['/nowhere'], clock=clock)
snap = hub.snapshot()
sub._store(snap, full=True)
updated = snap['state']['display']['last_updated']
sub._store({'version': 1, 'epoch': 'e1',
'volatile': {'display': {'new_key': 1}, 'plugins': {'published_at': 5.0},
'on_demand': 'junk'}}, full=False)
state = sub.latest()['state']
assert 'new_key' not in state['display'] and state['plugins'] is None
assert state['on_demand'] is None
# A tick for a version this copy is not at says nothing about it.
sub._store({'version': 7, 'epoch': 'e1',
'volatile': {'display': {'last_updated': 5.0}}}, full=False)
assert sub.latest()['state']['display']['last_updated'] == updated
# A display from before ticks carried them: the loop still refreshes.
sub._store({'version': 1, 'epoch': 'e1', 'loop': {'heartbeat_age_seconds': 3.0}},
full=False)
assert sub.latest()['loop'] == {'heartbeat_age_seconds': 3.0}
def test_a_tick_from_another_epoch_is_ignored(self, hub):
clock = FakeClock()
sub = client.StateSubscription(paths=['/nowhere'], clock=clock)
sub._store(hub.snapshot(), full=True)
clock.now += 10
sub._store({'version': 9, 'epoch': 'other', 'loop': {}}, full=False)
assert sub.latest()['received_mono'] == clock.now - 10
def test_loop_age_counts_the_time_since_it_arrived(self, hub):
snap = dict(hub.snapshot(), received_mono=100.0)
assert client.snapshot_loop_age(snap, now_mono=104.0) == 5.0
snap['loop'] = {'heartbeat_age_seconds': None}
assert client.snapshot_loop_age(snap, now_mono=104.0) is None
class TestReconnectBackoff:
"""StateSubscription._run's waits between connections, without a socket."""
def test_a_connection_that_got_a_snapshot_starts_the_backoff_over(self, hub,
monkeypatch):
"""Three failed tries, then the display is back twice, restarting
each time, then gone again. Each restart is retried after the
shortest wait, not after whatever the waits had grown to."""
sub = client.StateSubscription(paths=['/nowhere'])
script = ['refused', 'refused', 'refused', 'snapshot', 'snapshot', 'refused']
waits = []
def follow():
step = script.pop(0)
if step == 'snapshot': # subscribed, then the display restarted
sub._store(hub.snapshot(), full=True)
raise client.ControlError('closed', 'the display closed the connection')
raise client.ControlError(step)
def wait(seconds):
waits.append(seconds)
return not script # True ends _run, as stop() would
monkeypatch.setattr(sub, '_follow', follow)
monkeypatch.setattr(sub._stop, 'wait', wait)
sub._run()
first = client._RECONNECT_MIN_SECONDS
assert waits == [first, 2 * first, 4 * first, first, first, 2 * first]
def test_a_display_without_the_stream_is_still_retried_slowly(self, monkeypatch):
sub = client.StateSubscription(paths=['/nowhere'])
waits = []
def follow():
raise client.ControlError('unknown_command')
def wait(seconds):
waits.append(seconds)
return len(waits) == 2
monkeypatch.setattr(sub, '_follow', follow)
monkeypatch.setattr(sub._stop, 'wait', wait)
sub._run()
assert waits == [client._RECONNECT_MAX_SECONDS] * 2
# --- a real socket ------------------------------------------------------------------
def _wait_until(predicate, timeout=5.0):
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
if predicate():
return True
time.sleep(0.01)
return predicate()
@needs_unix_sockets
class TestLiveStream:
@pytest.fixture
def path(self, tmp_path):
return str(tmp_path / 'control.sock')
@pytest.fixture
def live(self, path, hub):
server = ControlServer(path, state_hub=hub, keepalive=0.2, io_timeout=0.5,
max_clients=2, max_subscribers=3)
assert server.start()
yield server, hub
server.close()
def _subscribe(self, path):
sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
sock.settimeout(5)
sock.connect(path)
sock.sendall(c.encode_message({'v': 1, 'id': 's1', 'cmd': Command.STATE_SUBSCRIBE,
'args': {}}))
return sock, c.FrameReader()
def _next(self, sock, reader, pending):
while not pending:
data = sock.recv(65536)
assert data, 'the display hung up'
pending.extend(reader.feed(data))
return json.loads(pending.pop(0))
def test_state_get_over_the_socket(self, live, path):
snap = client.state_get(paths=[path])
assert snap['version'] == 1 and snap['state']['display']['mode'] == 'clock'
short = client.state_get(since=snap['version'], epoch=snap['epoch'], paths=[path])
assert short['changed'] is False
def test_subscribe_answers_then_pushes_changes_in_order(self, live, path):
_, hub = live
sock, reader = self._subscribe(path)
pending = []
first = self._next(sock, reader, pending)
assert first['ok'] is True and first['id'] == 's1' and first['result']['version'] == 1
hub.publish('display', _display(mode='weather'), volatile=('last_updated',))
versions = []
while True:
msg = self._next(sock, reader, pending)
assert c.is_event(msg) and msg['id'] == 's1'
if msg['event'] == 'state':
versions.append(msg['result']['version'])
assert msg['result']['state']['display']['mode'] == 'weather'
break
hub.publish('brightness', {'brightness': 30})
while True:
msg = self._next(sock, reader, pending)
if msg['event'] == 'state':
versions.append(msg['result']['version'])
break
assert versions == [2, 3]
sock.close()
def test_ticks_keep_a_quiet_subscription_alive(self, live, path):
sock, reader = self._subscribe(path)
pending = []
self._next(sock, reader, pending)
kinds = [self._next(sock, reader, pending)['event'] for _ in range(3)]
assert kinds == ['tick', 'tick', 'tick']
sock.close()
def test_ticks_carry_the_timestamps_the_version_ignores(self, live, path):
_, hub = live
sock, reader = self._subscribe(path)
pending = []
self._next(sock, reader, pending)
hub.publish('display', _display(updated=4242.0), volatile=('last_updated',))
for _ in range(3): # a tick already on its way may predate the publish
tick = self._next(sock, reader, pending)
assert tick['event'] == 'tick'
if tick['result']['volatile'] == {'display': {'last_updated': 4242.0}}:
break
else:
pytest.fail(f'no tick carried the new last_updated: {tick}')
sock.close()
def test_a_subscription_keeps_last_updated_current(self, live, path):
_, hub = live
sub = client.StateSubscription(paths=[path]).start()
try:
assert _wait_until(lambda: sub.latest() is not None)
hub.publish('display', _display(updated=4242.0), volatile=('last_updated',))
assert _wait_until(lambda: (sub.latest() or {}).get('state', {})
.get('display', {}).get('last_updated') == 4242.0)
assert sub.snapshots == 1 # ticks, not new snapshots
finally:
sub.stop()
def test_a_burst_is_coalesced_to_the_latest(self, live, path):
_, hub = live
sock, reader = self._subscribe(path)
pending = []
self._next(sock, reader, pending)
for i in range(200):
hub.publish('display', _display(mode=f'm{i}'), volatile=('last_updated',))
seen = []
while not seen or seen[-1] != hub.version:
msg = self._next(sock, reader, pending)
if msg['event'] == 'state':
seen.append(msg['result']['version'])
assert seen == sorted(seen) and len(seen) < 200
assert msg['result']['state']['display']['mode'] == 'm199'
sock.close()
def test_several_subscribers_and_commands_still_get_a_slot(self, live, path):
server, hub = live
subs = [client.StateSubscription(paths=[path]).start() for _ in range(3)]
try:
assert _wait_until(lambda: all(s.latest() for s in subs))
assert hub.subscribers == 3
# max_clients is 2 and three streams are open: subscribers gave
# their request slots back.
for _ in range(4):
assert client.ping(paths=[path]) == {'pong': True}
hub.publish('display', _display(mode='weather'), volatile=('last_updated',))
assert _wait_until(lambda: all(
s.latest()['state']['display']['mode'] == 'weather' for s in subs))
# A fourth is over the subscriber bound.
sock, reader = self._subscribe(path)
refused = self._next(sock, reader, [])
assert refused['ok'] is False and refused['error']['code'] == ErrorCode.BUSY
sock.close()
finally:
for s in subs:
s.stop()
assert _wait_until(lambda: hub.subscribers == 0)
def test_a_subscriber_that_never_reads_blocks_nobody(self, live, path):
"""The render thread keeps publishing at full speed, a reading
subscriber keeps up, and the stuck one is dropped."""
_, hub = live
stuck = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
stuck.setsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF, 4096)
stuck.connect(path)
stuck.sendall(c.encode_message({'v': 1, 'id': 'stuck',
'cmd': Command.STATE_SUBSCRIBE, 'args': {}}))
good = client.StateSubscription(paths=[path]).start()
try:
assert _wait_until(lambda: good.latest() is not None)
assert _wait_until(lambda: hub.subscribers == 2)
big = {f'p{i}': {'state': 'enabled', 'note': 'x' * 100} for i in range(100)}
worst = 0.0
deadline = time.monotonic() + 3.0
i = 0
while time.monotonic() < deadline:
i += 1
start = time.perf_counter()
hub.publish('plugins', {'schema': 1, 'n': i, 'plugins': big})
worst = max(worst, time.perf_counter() - start)
time.sleep(0.005)
assert worst < 0.05, f'a publish took {worst:.3f}s'
# io_timeout is 0.5 s: the stuck one is gone, the good one is not.
assert _wait_until(lambda: hub.subscribers == 1)
assert _wait_until(lambda: (good.latest() or {}).get('state', {})
.get('plugins', {}).get('n') == i)
finally:
stuck.close()
good.stop()
def test_close_ends_the_streams(self, path, hub):
server = ControlServer(path, state_hub=hub, keepalive=30.0)
assert server.start()
sub = client.StateSubscription(paths=[path]).start()
try:
assert _wait_until(lambda: sub.latest() is not None)
start = time.monotonic()
server.close()
assert _wait_until(lambda: hub.subscribers == 0, timeout=3.0)
assert time.monotonic() - start < 3.0
assert _wait_until(lambda: sub.latest() is None, timeout=3.0)
finally:
sub.stop()
def test_the_subscription_follows_a_restarted_display(self, path, hub):
server = ControlServer(path, state_hub=hub, keepalive=0.2)
assert server.start()
sub = client.StateSubscription(paths=[path]).start()
try:
assert _wait_until(lambda: sub.latest() is not None)
server.close()
assert _wait_until(lambda: sub.latest() is None)
hub2 = StateHub(loop_probe=_loop(), epoch='e2')
hub2.publish('display', _display(mode='restarted'), volatile=('last_updated',))
server2 = ControlServer(path, state_hub=hub2, keepalive=0.2)
assert server2.start()
try:
assert _wait_until(lambda: (sub.latest() or {}).get('epoch') == 'e2',
timeout=8.0)
assert sub.latest()['state']['display']['mode'] == 'restarted'
finally:
server2.close()
finally:
sub.stop()
def test_a_server_without_state_answers_an_error(self, path):
server = ControlServer(path) # no hub
assert server.start()
try:
with pytest.raises(client.ControlError) as err:
client.state_get(paths=[path])
assert err.value.reason == ErrorCode.INTERNAL
finally:
server.close()
def test_no_display_is_no_socket(self, path):
with pytest.raises(client.ControlError) as err:
client.state_get(paths=[path])
assert err.value.reason == 'no_socket'
+120
View File
@@ -0,0 +1,120 @@
"""src.common and src.plugin_system import their exports lazily (PEP 562).
The web interface imports both packages only for small submodules
(path_safety, snapshot_policy, store_manager, ...). When their __init__
imported every export eagerly, that dragged numpy, freetype and the plugin
manager into a process that never uses them -- about 13 MB of RSS on a Pi.
These tests pin both halves of the change: the package import stays light,
and every exported name still resolves to the very object its home module
defines, so ``isinstance`` and ``is`` checks behave as before.
"""
import importlib
import json
import subprocess
import sys
import textwrap
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parents[1]
#: Modules the bare package import must not load. numpy is the one that
#: matters for memory; the rest are what the eager __init__ used to import.
HEAVY = (
"numpy",
"freetype",
"src.adaptive_layout",
"src.common.api_helper",
"src.common.logo_helper",
"src.common.scroll_helper",
"src.plugin_system.base_plugin",
"src.plugin_system.plugin_manager",
)
def _run(script):
result = subprocess.run(
[sys.executable, "-c", textwrap.dedent(script)], cwd=str(REPO_ROOT),
capture_output=True, text=True, timeout=120)
assert result.returncode == 0, result.stderr
return json.loads(result.stdout.strip().splitlines()[-1])
@pytest.mark.parametrize("package", ["src.common", "src.plugin_system"])
def test_package_import_loads_nothing_heavy(package):
# A fresh interpreter: this test process has long since imported them.
loaded = _run(f"""
import json, sys
import {package}
print(json.dumps(sorted(m for m in {HEAVY!r} if m in sys.modules)))
""")
assert loaded == [], f"importing {package} loaded {loaded}"
def test_web_interface_submodules_do_not_load_numpy():
# The imports web_interface/app.py and its API blueprint make from these
# packages. numpy here costs the web process ~13 MB for nothing.
loaded = _run("""
import json, sys
from src.common import path_safety, snapshot_policy, sync_manager
from src.plugin_system import store_manager, schema_manager
print(json.dumps("numpy" in sys.modules))
""")
assert loaded is False
@pytest.mark.parametrize("package", ["src.common", "src.plugin_system"])
def test_every_exported_name_resolves_to_its_home_object(package):
pkg = importlib.import_module(package)
assert sorted(pkg.__all__) == sorted(pkg._LAZY), "__all__ and _LAZY differ"
for name in pkg.__all__:
module_name, attr = pkg._LAZY[name]
home = importlib.import_module(module_name)
expected = home if attr is None else getattr(home, attr)
assert getattr(pkg, name) is expected, name
assert name in dir(pkg)
def test_from_import_forms_plugins_use():
# Every form found in ledmatrix-plugins and core: names, aliases,
# submodules through the package, dotted submodule imports.
from src.common import ScrollHelper, LogoHelper
from src.common import scroll_config as _scroll_config
from src.common import sports_card as _card
from src.common import draw_fitted_text
from src.plugin_system import BasePlugin, PluginManager
from src.plugin_system import compatibility
import src.common.scroll_helper
import src.plugin_system.base_plugin
import src.common
import src.plugin_system
assert ScrollHelper is src.common.scroll_helper.ScrollHelper
assert LogoHelper is src.common.LogoHelper
assert _scroll_config is src.common.scroll_config
assert _card is importlib.import_module("src.common.sports_card")
assert draw_fitted_text is importlib.import_module("src.adaptive_layout").draw_fitted_text
assert BasePlugin is src.plugin_system.base_plugin.BasePlugin
assert PluginManager is importlib.import_module(
"src.plugin_system.plugin_manager").PluginManager
assert compatibility is importlib.import_module("src.plugin_system.compatibility")
assert src.plugin_system.__version__ == "1.0.0"
def test_star_import_still_binds_everything():
namespace = {}
exec("from src.common import *", namespace)
import src.common
assert set(src.common.__all__) <= set(namespace)
@pytest.mark.parametrize("package", ["src.common", "src.plugin_system"])
def test_unknown_name_raises_attribute_error(package):
pkg = importlib.import_module(package)
with pytest.raises(AttributeError, match="no_such_name"):
pkg.no_such_name # noqa: B018
with pytest.raises(ImportError):
exec(f"from {package} import no_such_name")
@@ -58,3 +58,83 @@ def test_a_reloaded_plugin_still_gets_its_own_bare_module(plugins):
assert reloaded.WHO == "alpha" assert reloaded.WHO == "alpha"
assert sys.path.index(str(plugins["alpha"])) < sys.path.index(str(plugins["beta"])) assert sys.path.index(str(plugins["alpha"])) < sys.path.index(str(plugins["beta"]))
assert sys.path.count(str(plugins["alpha"])) == 1 assert sys.path.count(str(plugins["alpha"])) == 1
# -- sub-packages ------------------------------------------------------------
#
# A plugin that keeps helpers in a package (``providers/feed.py``, imported as
# ``from providers.feed import ...``) leaves dotted entries in sys.modules.
# Only the bare ``providers`` used to be tracked, so ``providers.feed`` outlived
# the plugin: a reload after a store update re-ran the new manager.py against
# the old feed.py, until the display restarted. Elections (providers/),
# flights (enrichment/) and olympics (data/, renderers/) ship packages.
@pytest.fixture
def package_plugin(tmp_path):
before_path = list(sys.path)
before_modules = set(sys.modules)
plugin_dir = tmp_path / "pkgdemo"
(plugin_dir / "providers").mkdir(parents=True)
(plugin_dir / "providers" / "__init__.py").write_text("", encoding="utf-8")
(plugin_dir / "providers" / "feed.py").write_text("VERSION = 'v1'\n", encoding="utf-8")
(plugin_dir / "manager.py").write_text(
"from providers.feed import VERSION\n", encoding="utf-8")
yield plugin_dir
sys.path[:] = before_path
for key in set(sys.modules) - before_modules:
sys.modules.pop(key, None)
def test_a_reloaded_plugin_runs_its_updated_subpackage_module(package_plugin):
loader = PluginLoader()
assert loader.load_module("pkgdemo", package_plugin, "manager.py").VERSION == "v1"
_unload(loader, "pkgdemo")
# The store update: a different size, so no cached bytecode can match.
(package_plugin / "providers" / "feed.py").write_text(
"VERSION = 'v2 from the update'\n", encoding="utf-8")
reloaded = loader.load_module("pkgdemo", package_plugin, "manager.py")
assert reloaded.VERSION == "v2 from the update"
def test_unload_drops_the_plugins_subpackage_modules(package_plugin):
loader = PluginLoader()
loader.load_module("pkgdemo", package_plugin, "manager.py")
# Still importable while the plugin runs, as before.
assert "providers.feed" in sys.modules
_unload(loader, "pkgdemo")
assert not [k for k in sys.modules if k.startswith("providers")]
def test_a_failed_load_leaves_no_subpackage_module_behind(package_plugin):
(package_plugin / "manager.py").write_text(
"from providers.feed import VERSION\nraise RuntimeError('broken')\n",
encoding="utf-8")
loader = PluginLoader()
with pytest.raises(RuntimeError):
loader.load_module("pkgdemo", package_plugin, "manager.py")
assert not [k for k in sys.modules if k.startswith("providers")]
def test_unload_leaves_packages_from_outside_the_plugin_alone(package_plugin, tmp_path):
# A library the plugin imports is not the plugin's to drop.
lib_root = tmp_path / "site"
(lib_root / "extlib").mkdir(parents=True)
(lib_root / "extlib" / "__init__.py").write_text("", encoding="utf-8")
(lib_root / "extlib" / "sub.py").write_text("X = 1\n", encoding="utf-8")
sys.path.append(str(lib_root))
(package_plugin / "manager.py").write_text(
"import extlib.sub\nfrom providers.feed import VERSION\n", encoding="utf-8")
loader = PluginLoader()
loader.load_module("pkgdemo", package_plugin, "manager.py")
_unload(loader, "pkgdemo")
assert "extlib.sub" in sys.modules
assert "extlib" in sys.modules
+81
View File
@@ -0,0 +1,81 @@
"""A dev plugin linked in under a name its checkout does not share still loads.
``scripts/dev/dev_plugin_setup.sh`` links a checkout into the plugins
directory under the plugin's id: ``link-github foo <url>`` clones
``ledmatrix-foo`` (the repository naming convention) and links it as
``plugins/foo``. ``contained_plugin_dir`` resolved the link and looked for the
*target's* folder name, ``ledmatrix-foo``, among the plugins directory's
entries. There is none, so ``install_dependencies`` refused the plugin as
outside the plugins directory and the load failed with "Dependency
installation failed" -- even with no requirements.txt at all.
The containment it exists for still holds: the answer is always rebuilt from
an entry enumerated under the plugins directory.
Skipped where this process cannot create a symlink (Windows without the
privilege).
"""
import os
from unittest.mock import MagicMock, patch
import pytest
from src.plugin_system.plugin_loader import PluginLoader, contained_plugin_dir
def _symlink_or_skip(target, link):
try:
os.symlink(target, link, target_is_directory=True)
except (OSError, NotImplementedError) as e:
pytest.skip(f"cannot create a symlink here: {e}")
@pytest.fixture
def linked(tmp_path):
checkout = tmp_path / "dev-plugins" / "ledmatrix-foo"
checkout.mkdir(parents=True)
plugins_dir = tmp_path / "plugins"
plugins_dir.mkdir()
link = plugins_dir / "foo"
_symlink_or_skip(checkout, link)
return plugins_dir, link, checkout
def test_a_link_resolves_to_its_own_entry_in_the_plugins_dir(linked):
plugins_dir, link, _checkout = linked
assert contained_plugin_dir(link, plugins_dir) == os.path.join(
os.path.realpath(plugins_dir), "foo")
def test_a_linked_plugin_without_requirements_needs_no_install(linked):
plugins_dir, link, _checkout = linked
with patch("subprocess.run") as pip:
assert PluginLoader().install_dependencies(link, "foo", plugins_dir=plugins_dir) is True
pip.assert_not_called()
@patch("src.plugin_system.plugin_loader.requirements_are_satisfied", return_value=False)
def test_a_linked_plugins_requirements_are_installed_through_the_link(_satisfied, linked):
plugins_dir, link, checkout = linked
(checkout / "requirements.txt").write_text("package1==1.0.0\n", encoding="utf-8")
with patch("subprocess.run", return_value=MagicMock(returncode=0, stderr="")) as pip:
assert PluginLoader().install_dependencies(link, "foo", plugins_dir=plugins_dir) is True
argv = pip.call_args[0][0]
assert argv[argv.index("-r") + 1] == os.path.join(
os.path.realpath(plugins_dir), "foo", "requirements.txt")
def test_a_link_outside_the_plugins_dir_is_still_refused(linked, tmp_path):
plugins_dir, _link, checkout = linked
elsewhere = tmp_path / "elsewhere"
elsewhere.mkdir()
stray = elsewhere / "bar"
_symlink_or_skip(checkout, stray)
assert contained_plugin_dir(stray, plugins_dir) is None
assert contained_plugin_dir(plugins_dir / ".." / "elsewhere" / "bar", plugins_dir) is None
+18
View File
@@ -189,6 +189,24 @@ class TestPluginExecutor:
assert result is False assert result is False
def test_a_base_exception_is_a_failure_not_a_timeout(self):
"""asyncio.CancelledError derives from BaseException. Uncaught on
the executor's thread it ended the thread with the call never marked
complete, so a call that failed at once was reported, and recorded,
as timing out."""
import asyncio
import pytest
from src.exceptions import PluginError
from src.plugin_system.plugin_executor import PluginExecutor
executor = PluginExecutor(default_timeout=5.0)
def cancelled():
raise asyncio.CancelledError()
with pytest.raises(PluginError) as raised:
executor.execute_with_timeout(cancelled, plugin_id="test_plugin")
assert isinstance(raised.value.__cause__, asyncio.CancelledError)
class TestPluginHealth: class TestPluginHealth:
"""Test plugin health monitoring.""" """Test plugin health monitoring."""
+134
View File
@@ -0,0 +1,134 @@
"""The render loop's scheduled-update pass, throttled per frame.
The frame loops called _tick_plugin_updates() after every frame, about 125
times a second on a scroller, and each call ran
PluginManager.run_scheduled_updates(): a copy of the plugin dict and several
locks per plugin, to find that nothing was due (no interval is shorter than
5 s). The frame loops and the dwell sleep now call
_tick_plugin_updates_if_due(), which runs it at most once per
PLUGIN_UPDATE_TICK_INTERVAL. The top of each loop pass still calls
_tick_plugin_updates() unthrottled, because a plugin loaded, reloaded or
enabled for on-demand there is due at once.
"""
import os
from types import SimpleNamespace
from unittest.mock import MagicMock
import pytest
os.environ.setdefault("EMULATOR", "true")
from src import display_controller as dc_mod # noqa: E402
from src.display_controller import DisplayController # noqa: E402
from test._run_loop_harness import FakePlugin, RunLoopHarness # noqa: E402
@pytest.fixture
def clock(monkeypatch):
now = {"t": 1000.0}
fake = SimpleNamespace(monotonic=lambda: now["t"], time=lambda: now["t"])
monkeypatch.setattr(dc_mod, "time", fake)
return now
@pytest.fixture
def dc():
controller = object.__new__(DisplayController)
controller.plugin_manager = MagicMock()
return controller
def _calls(dc):
return dc.plugin_manager.run_scheduled_updates.call_count
class TestThrottle:
def test_the_first_call_runs(self, dc, clock):
dc._tick_plugin_updates_if_due()
assert _calls(dc) == 1
def test_calls_inside_the_window_are_skipped(self, dc, clock):
dc._tick_plugin_updates_if_due()
for _ in range(30): # a scroller's frames, 8 ms apart
clock["t"] += 0.008
dc._tick_plugin_updates_if_due()
assert clock["t"] - 1000.0 < dc.PLUGIN_UPDATE_TICK_INTERVAL
assert _calls(dc) == 1
def test_it_runs_again_once_the_window_has_passed(self, dc, clock):
dc._tick_plugin_updates_if_due()
clock["t"] += dc.PLUGIN_UPDATE_TICK_INTERVAL
dc._tick_plugin_updates_if_due()
assert _calls(dc) == 2
def test_the_unthrottled_tick_always_runs_and_restarts_the_window(self, dc, clock):
# The top of the loop pass: a plugin just (re)loaded is due now.
dc._tick_plugin_updates_if_due()
clock["t"] += 0.01
dc._tick_plugin_updates()
assert _calls(dc) == 2
# ...and the frame right after it does not repeat the pass.
clock["t"] += 0.01
dc._tick_plugin_updates_if_due()
assert _calls(dc) == 2
def test_no_plugin_manager(self, clock):
controller = object.__new__(DisplayController)
controller.plugin_manager = None
controller._tick_plugin_updates_if_due() # must not raise
controller._tick_plugin_updates()
def test_a_failing_pass_is_contained_and_still_throttled(self, dc, clock):
dc.plugin_manager.run_scheduled_updates.side_effect = RuntimeError("boom")
dc._tick_plugin_updates_if_due()
dc._tick_plugin_updates_if_due()
assert _calls(dc) == 1
def test_the_vegas_tick_is_not_throttled(self, dc, clock):
# Vegas's own update thread calls the plugin manager directly.
dc.vegas_coordinator = None
dc.plugin_manager.run_scheduled_updates_with_changes.return_value = []
dc._tick_plugin_updates()
dc._tick_plugin_updates_for_vegas()
dc._tick_plugin_updates_for_vegas()
assert dc.plugin_manager.run_scheduled_updates_with_changes.call_count == 2
class TestInTheRunLoop:
"""The real run() on the golden-trace harness's fake clock."""
@pytest.fixture
def run(self, tmp_path):
h = RunLoopHarness(tmp_path, horizon=40)
h.add_plugin(FakePlugin("ticker", ["ticker"], duration=10, enable_scrolling=True))
h.add_plugin(FakePlugin("clock", ["clock"], duration=10))
ticks = []
h.pm.run_scheduled_updates = lambda: ticks.append(round(h.clock.rel(), 3))
h.run()
passes = sorted({t for t, kind, _s, _d in h.events if kind == "pass"})
frames = [t for t, kind, s, _d in h.events if kind in ("first", "frame") and s == "ticker"]
return SimpleNamespace(ticks=ticks, passes=passes, frames=frames,
interval=h.controller.PLUGIN_UPDATE_TICK_INTERVAL)
def test_a_scroller_ticks_a_few_times_a_second_not_every_frame(self, run):
ticker_frames = len(run.frames)
assert ticker_frames > 500 # 10 s screens at 125 Hz, twice
# 40 s at four a second, plus one per pass.
assert len(run.ticks) <= 40 / run.interval + len(run.passes) + 1
def test_ticks_are_never_closer_than_the_floor_except_at_a_pass(self, run):
passes = set(run.passes)
for prev, cur in zip(run.ticks, run.ticks[1:]):
if cur - prev < run.interval - 1e-6:
assert cur in passes, (prev, cur)
def test_every_pass_ticks_at_once(self, run):
# Unthrottled: the frame loop ticked just before the screen ended, and
# the pass ticks again anyway, so a plugin it just loaded is not kept
# waiting.
ticks = set(run.ticks)
for t in run.passes:
assert t in ticks, t
assert any(cur - prev < run.interval
for prev, cur in zip(run.ticks, run.ticks[1:]))
+525
View File
@@ -0,0 +1,525 @@
"""Updates refresh the installed systemd units: the root helper and its callers.
An update moved the checkout, and with it systemd/*.service, but systemd runs
the copies in /etc/systemd/system, which only the installer wrote. So unit
settings added after a device was installed (#687's render-loop watchdog)
never reached it. scripts/install/ledmatrix_refresh_units.py, installed
root-owned as /usr/local/sbin/ledmatrix-refresh-units and granted to the web
user by exact command line, now installs changed units after an update, and
puts the previous ones back when the automatic update rolls back.
The helper runs as root on input the web user can edit (the templates), so
most of these are about what it refuses.
"""
import importlib.util
import json
import os
import re
import shutil
import subprocess
import sys
from pathlib import Path
import pytest
ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(ROOT))
_spec = importlib.util.spec_from_file_location(
'ledmatrix_refresh_units', ROOT / 'scripts' / 'install' / 'ledmatrix_refresh_units.py')
ru = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(ru)
from web_interface import unit_refresh # noqa: E402
try:
import pwd
# The web side checks the account exists, so use one that does.
WEB_USER = pwd.getpwuid(os.getuid()).pw_name
except ImportError: # Windows
WEB_USER = 'ledpi'
# An account a hostile template switches the web unit to. Root, unless the
# tests themselves run as root: then root *is* the web user and switching to
# it changes nothing, so use another account.
OTHER_USER = 'root' if WEB_USER != 'root' else 'nobody'
class Host:
"""A project checkout, an /etc/systemd/system, and a fake systemctl."""
def __init__(self, tmp_path, web_user=WEB_USER):
self.project = tmp_path / 'LEDMatrix'
shutil.copytree(ROOT / 'systemd', self.project / 'systemd')
self.systemd = tmp_path / 'etc-systemd-system'
self.systemd.mkdir()
self.backup = tmp_path / 'var-lib-ledmatrix' / 'unit-backup'
self.web_user = web_user
self.calls = []
self.path_active = True
self.root = True
for name in ru.UNITS:
self.install(name)
def template(self, name):
return (self.project / 'systemd' / name).read_text(encoding='utf-8')
def set_template(self, name, text):
(self.project / 'systemd' / name).write_text(text, encoding='utf-8', newline='\n')
def rendered(self, name, text=None):
user = 'root' if name == ru.DISPLAY_UNIT else self.web_user
return ru.render(text if text is not None else self.template(name), str(self.project), user)
def install(self, name, text=None):
(self.systemd / name).write_text(self.rendered(name, text), encoding='utf-8', newline='\n')
def installed(self, name):
path = self.systemd / name
return path.read_text(encoding='utf-8') if path.exists() else None
def run(self, args, **kwargs):
self.calls.append(list(args))
if args[:2] == ['systemctl', 'is-active']:
out = 'active\n' if self.path_active else 'inactive\n'
return subprocess.CompletedProcess(args, 0, stdout=out, stderr='')
return subprocess.CompletedProcess(args, 0, stdout='', stderr='')
def refresher(self, log=None):
return ru.Refresher(systemd_dir=str(self.systemd), backup_dir=str(self.backup), run=self.run,
is_root=lambda: self.root, user_exists=lambda user: True,
log=log or (lambda m: None))
@property
def reloads(self):
return self.calls.count(['systemctl', 'daemon-reload'])
def watchdog_added(text):
"""The kind of change #687 made: a new directive in [Service]."""
return text.replace('[Service]\n', '[Service]\nWatchdogSec=60\n', 1)
@pytest.fixture
def host(tmp_path):
return Host(tmp_path)
# -- refresh ------------------------------------------------------------------
def test_units_that_match_are_left_alone(host):
assert host.refresher().refresh() == []
assert host.reloads == 0
def test_a_changed_template_is_installed_and_systemd_reloaded(host):
old = host.installed(ru.DISPLAY_UNIT)
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
assert host.refresher().refresh() == [ru.DISPLAY_UNIT]
assert 'WatchdogSec=60' in host.installed(ru.DISPLAY_UNIT)
assert host.installed(ru.DISPLAY_UNIT) == host.rendered(ru.DISPLAY_UNIT)
assert host.reloads == 1
# Only the unit that changed is replaced, and the one it replaced is kept.
assert (host.backup / ru.DISPLAY_UNIT).read_text(encoding='utf-8') == old
assert json.loads((host.backup / ru.MANIFEST).read_text()) == {'units': [ru.DISPLAY_UNIT]}
def test_the_web_unit_keeps_the_web_users_account(host):
host.set_template(ru.WEB_UNIT, watchdog_added(host.template(ru.WEB_UNIT)))
host.refresher().refresh()
assert ru.directive_values(host.installed(ru.WEB_UNIT), 'User') == [WEB_USER]
assert ru.directive_values(host.installed(ru.DISPLAY_UNIT), 'User') == ['root']
def test_comment_only_changes_are_not_a_refresh(host):
host.set_template(ru.DISPLAY_UNIT, '# a new comment\n\n' + host.template(ru.DISPLAY_UNIT))
assert host.refresher().refresh() == []
assert host.reloads == 0
def test_a_unit_that_was_never_installed_is_not_installed(host):
(host.systemd / ru.VERIFY_SERVICE).unlink()
(host.systemd / ru.VERIFY_PATH).unlink()
host.set_template(ru.VERIFY_SERVICE, watchdog_added(host.template(ru.VERIFY_SERVICE)))
host.refresher().refresh()
assert host.installed(ru.VERIFY_SERVICE) is None
def test_a_changed_path_unit_is_restarted_so_it_watches_the_new_path(host):
host.set_template(ru.VERIFY_PATH, host.template(ru.VERIFY_PATH).replace(
'[Path]\n', '[Path]\nMakeDirectory=yes\n'))
host.refresher().refresh()
assert ['systemctl', 'restart', ru.VERIFY_PATH] in host.calls
def test_it_must_run_as_root(host):
host.root = False
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
with pytest.raises(ru.RefreshError, match='root'):
host.refresher().refresh()
assert 'WatchdogSec=60' not in host.installed(ru.DISPLAY_UNIT)
def _failing_reload(host):
real = host.run
def run(args, **kwargs):
if args == ['systemctl', 'daemon-reload'] and host.reloads == 0:
host.calls.append(list(args))
return subprocess.CompletedProcess(args, 1, stdout='', stderr='boom')
return real(args, **kwargs)
return run
def test_a_failed_daemon_reload_puts_the_old_units_back(host):
# The web side reports this as a failure and records no units_refreshed,
# so a rollback would not --restore: the helper must undo it itself.
old = host.installed(ru.DISPLAY_UNIT)
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
host.run = _failing_reload(host)
with pytest.raises(ru.RefreshError):
host.refresher().refresh()
assert host.installed(ru.DISPLAY_UNIT) == old
assert host.reloads == 2 # the failed one, then one after putting it back
assert not (host.backup / ru.MANIFEST).exists()
def test_a_failed_write_puts_back_the_units_already_written(host, monkeypatch):
old = {name: host.installed(name) for name in (ru.DISPLAY_UNIT, ru.WEB_UNIT)}
for name in old:
host.set_template(name, watchdog_added(host.template(name)))
refresher = host.refresher()
real_write = refresher._write_unit
writes = []
def write(name, text):
writes.append(name)
if len(writes) == 2:
raise OSError('disk full')
real_write(name, text)
monkeypatch.setattr(refresher, '_write_unit', write)
with pytest.raises(OSError):
refresher.refresh()
first = writes[0]
assert host.installed(first) == old[first]
assert all(host.installed(name) == old[name] for name in old)
# -- what it refuses ------------------------------------------------------------
@pytest.mark.parametrize('unit, edit', [
# The web interface's unit switched to root by a template edit.
(ru.WEB_UNIT, lambda t: t.replace('User=__USER__', f'User={OTHER_USER}')),
# The display's unit switched to another account.
(ru.DISPLAY_UNIT, lambda t: t.replace('User=root', 'User=nobody')),
# A second User= line.
(ru.VERIFY_SERVICE, lambda t: t.replace('[Service]\n', '[Service]\nUser=root\n', 1)),
# Run from somewhere else.
(ru.DISPLAY_UNIT, lambda t: t.replace('WorkingDirectory=__PROJECT_ROOT_DIR__', 'WorkingDirectory=/tmp')),
# A second User= written with spaces, which systemd accepts (last one wins).
# Placed in [Service] (before [Install]), where it is not a layout problem.
(ru.WEB_UNIT, lambda t: t.replace('\n[Install]', 'User = root\n\n[Install]')),
# The web user's User= moved to [Unit], where systemd ignores it (so root).
(ru.WEB_UNIT, lambda t: t.replace('User=__USER__\n', '').replace('[Unit]\n', '[Unit]\nUser=__USER__\n')),
# The User= line hidden inside a continued line, where systemd does not see it.
(ru.WEB_UNIT, lambda t: t.replace('User=__USER__\n', '').replace(
'Description=LED Matrix Web Interface Service\n',
'Description=LED Matrix Web Interface Service \\\\\nUser=__USER__\n')),
# A path unit that starts something else.
(ru.VERIFY_PATH, lambda t: t.replace('Unit=ledmatrix-update-verify.service', 'Unit=ledmatrix.service')),
])
def test_a_template_that_changes_who_or_where_is_refused_and_nothing_changes(host, unit, edit):
before = {name: host.installed(name) for name in ru.UNITS}
# A legitimate change alongside, which must not go in either.
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
host.set_template(unit, edit(host.template(unit)))
with pytest.raises(ru.RefreshError):
host.refresher().refresh()
assert {name: host.installed(name) for name in ru.UNITS} == before
assert host.reloads == 0
def test_the_project_folder_comes_from_the_installed_unit_not_the_caller(host, tmp_path):
# The installed display unit names the project; a WorkingDirectory that
# is not an existing absolute folder is refused before any template is read.
host.install(ru.DISPLAY_UNIT, host.template(ru.DISPLAY_UNIT).replace(
'WorkingDirectory=__PROJECT_ROOT_DIR__', 'WorkingDirectory=relative/path'))
with pytest.raises(ru.RefreshError, match='cannot be used'):
host.refresher().plan()
def test_without_the_display_unit_installed_nothing_is_done(host):
(host.systemd / ru.DISPLAY_UNIT).unlink()
with pytest.raises(ru.RefreshError, match='not installed'):
host.refresher().plan()
@pytest.mark.skipif(not hasattr(os, 'O_NOFOLLOW'), reason='POSIX only')
def test_a_template_symlink_is_not_followed(host, tmp_path):
secret = tmp_path / 'secret'
secret.write_text(host.template(ru.DISPLAY_UNIT) + 'Environment=SECRET=1\n', encoding='utf-8')
target = host.project / 'systemd' / ru.DISPLAY_UNIT
target.unlink()
target.symlink_to(secret)
with pytest.raises(ru.RefreshError):
host.refresher().plan()
@pytest.mark.skipif(not hasattr(os, 'O_NOFOLLOW'), reason='POSIX only')
def test_a_symlinked_systemd_folder_is_not_followed(host, tmp_path):
elsewhere = tmp_path / 'elsewhere'
shutil.move(str(host.project / 'systemd'), str(elsewhere))
(host.project / 'systemd').symlink_to(elsewhere, target_is_directory=True)
with pytest.raises(ru.RefreshError):
host.refresher().plan()
def test_an_oversized_template_is_refused(host):
host.set_template(ru.DISPLAY_UNIT, host.template(ru.DISPLAY_UNIT) + '#' * (ru.MAX_TEMPLATE_BYTES + 1))
with pytest.raises(ru.RefreshError, match='larger'):
host.refresher().plan()
def test_main_leaves_the_callers_environment_alone(host):
"""main() runs in-process in these tests; pinning PATH belongs to the installed program."""
before = os.environ.get('PATH')
ru.main(['ledmatrix-refresh-units', '--check'], refresher=host.refresher())
assert os.environ.get('PATH') == before
@pytest.mark.parametrize('argv', [
['--restore', 'x'], ['--refresh'], ['/etc/passwd'], ['--check', '--restore'], ['']])
def test_any_other_command_line_is_refused(argv):
class Boom:
def __getattr__(self, name):
raise AssertionError('must not run')
assert ru.main(['ledmatrix-refresh-units', *argv], refresher=Boom()) == ru.EXIT_USAGE
def test_main_reports_a_refusal_as_a_failure(host, capsys):
host.set_template(ru.WEB_UNIT, host.template(ru.WEB_UNIT).replace('User=__USER__', f'User={OTHER_USER}'))
assert ru.main(['ledmatrix-refresh-units'], refresher=host.refresher()) == ru.EXIT_FAILED
assert 'refusing' in capsys.readouterr().err
# -- restore ------------------------------------------------------------------
def test_restore_puts_back_exactly_what_the_refresh_replaced(host):
# A hand-edited installed unit: the rollback must give back this file,
# not a rendering of the old template.
hand_edited = host.installed(ru.DISPLAY_UNIT) + '# edited by hand\n'
(host.systemd / ru.DISPLAY_UNIT).write_text(hand_edited, encoding='utf-8', newline='\n')
untouched = host.installed(ru.WEB_UNIT)
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
host.refresher().refresh()
assert host.refresher().restore() == [ru.DISPLAY_UNIT]
assert host.installed(ru.DISPLAY_UNIT) == hand_edited
assert host.installed(ru.WEB_UNIT) == untouched
assert host.reloads == 2
assert not (host.backup / ru.MANIFEST).exists(), 'a second restore must not repeat it'
assert host.refresher().restore() == []
def test_restore_after_an_update_that_changed_no_units_restores_nothing(host):
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
host.refresher().refresh() # an earlier update...
newer = host.installed(ru.DISPLAY_UNIT)
host.refresher().refresh() # ...then one that changed no units
assert host.refresher().restore() == []
assert host.installed(ru.DISPLAY_UNIT) == newer
def test_restore_must_run_as_root(host):
host.root = False
with pytest.raises(ru.RefreshError, match='root'):
host.refresher().restore()
# -- the real templates and install_service.sh ----------------------------------
def test_every_shipped_template_passes_the_helpers_checks(tmp_path):
host = Host(tmp_path)
for name in ru.UNITS:
host.set_template(name, watchdog_added(host.template(name)) if name.endswith('.service')
else host.template(name))
assert host.refresher().refresh() == sorted(n for n in ru.UNITS if n.endswith('.service'))
@pytest.mark.skipif(not sys.platform.startswith('linux'), reason='runs sed as install_service.sh does')
def test_rendering_matches_install_service_sh(tmp_path):
"""Same text as the installer's sed, including characters sed treats specially."""
lib = ROOT / 'scripts' / 'install' / 'lib_systemd_render.sh'
for project in ('/home/pi/LEDMatrix', '/opt/led matrix&co'):
for name in ru.UNITS:
user = 'root' if name == ru.DISPLAY_UNIT else 'pi'
script = (f'source "{lib}"; R=$(sed_escape_replacement "$1"); U=$(sed_escape_replacement "$2"); '
f'sed "s|__PROJECT_ROOT_DIR__|$R|g; s|__USER__|$U|g" "$3"')
out = subprocess.run(['bash', '-c', script, 'x', project, user, str(ROOT / 'systemd' / name)],
capture_output=True, text=True, check=True).stdout
template = (ROOT / 'systemd' / name).read_text(encoding='utf-8')
assert ru.render(template, project, user) == out, name
def test_install_service_installs_the_helper_root_owned_at_the_granted_path():
text = (ROOT / 'scripts' / 'install' / 'install_service.sh').read_text(encoding='utf-8')
assert 'scripts/install/ledmatrix_refresh_units.py' in text
m = re.search(r'install -D -o root -g root -m 0755 "\$REFRESH_UNITS_SRC" "\$REFRESH_UNITS_DEST"', text)
assert m, 'install_service.sh must install the helper root:root 0755'
assert f'REFRESH_UNITS_DEST={ru.INSTALLED_PATH}' in text
def test_every_caller_names_the_same_helper_path():
lib = (ROOT / 'scripts' / 'install' / 'lib_sudoers.sh').read_text(encoding='utf-8')
verifier = (ROOT / 'scripts' / 'utils' / 'auto_update_verify.py').read_text(encoding='utf-8')
assert f'LEDMATRIX_REFRESH_UNITS_PATH={ru.INSTALLED_PATH}' in lib
assert unit_refresh.HELPER_PATH == ru.INSTALLED_PATH
assert f"REFRESH_UNITS_PATH = '{ru.INSTALLED_PATH}'" in verifier
assert ru.INSTALLED_PATH.startswith('/usr/local/sbin/'), 'must live outside the user-owned checkout'
def test_the_helper_imports_nothing_from_the_checkout():
source = (ROOT / 'scripts' / 'install' / 'ledmatrix_refresh_units.py').read_text(encoding='utf-8')
imports = re.findall(r'^\s*(?:from|import)\s+([\w.]+)', source, re.M)
assert not [m for m in imports if m.split('.')[0] in ('src', 'web_interface', 'scripts')]
assert source.startswith('#!/usr/bin/python3 -I\n'), 'isolated mode: no PYTHON* env, no user site'
# -- the web interface's side (web_interface/unit_refresh.py) ---------------------
class Sudo:
"""sudo: refuses (``rc``/``stderr``), or runs the real helper as root against ``host``."""
def __init__(self, host=None, rc=0, stderr=''):
self.host, self.rc, self.stderr, self.calls = host, rc, stderr, []
self.as_root = False
def __call__(self, args, **kwargs):
self.calls.append(list(args))
if self.rc or self.host is None:
return subprocess.CompletedProcess(args, self.rc, stdout='', stderr=self.stderr)
lines = []
self.as_root = True
try:
rc = ru.main(['ledmatrix-refresh-units', *args[3:]], refresher=self.host.refresher(lines.append))
finally:
self.as_root = False
return subprocess.CompletedProcess(args, rc, stdout='\n'.join(lines) + '\n', stderr='')
def _web(host, tmp_path, sudo, helper_installed=True):
helper = tmp_path / 'usr-local-sbin' / 'ledmatrix-refresh-units'
if helper_installed:
helper.parent.mkdir(exist_ok=True)
helper.write_text('#!/bin/true\n')
return unit_refresh.refresh_after_update(run=sudo, systemd_dir=str(host.systemd),
helper_path=str(helper))
def _stale(host):
# The web side compares the installed units with the checkout's templates.
host.set_template(ru.DISPLAY_UNIT, watchdog_added(host.template(ru.DISPLAY_UNIT)))
def test_web_side_does_nothing_when_the_units_match(host, tmp_path):
sudo = Sudo(host)
result = _web(host, tmp_path, sudo)
assert result['status'] == unit_refresh.CURRENT and sudo.calls == []
assert result['message'] == ''
def test_web_side_runs_the_helper_through_sudo_with_no_arguments(host, tmp_path):
_stale(host)
sudo = Sudo(host)
result = _web(host, tmp_path, sudo)
assert result['status'] == unit_refresh.REFRESHED
assert result['units'] == [ru.DISPLAY_UNIT]
assert len(sudo.calls) == 1 and sudo.calls[0][:2] == ['sudo', '-n'] and len(sudo.calls[0]) == 3
assert 'ledmatrix.service' in result['message']
assert 'WatchdogSec=60' in host.installed(ru.DISPLAY_UNIT)
@pytest.fixture
def unreadable(monkeypatch, host):
"""Installed units only root can read (install_service.sh used to leave them 0600)."""
real = ru._read_installed
sudo = Sudo(host)
def read(systemd_dir, name):
if not sudo.as_root:
raise ru.UnitsUnreadable(f'cannot read the installed {name}: Permission denied')
return real(systemd_dir, name)
monkeypatch.setattr(ru, '_read_installed', read)
monkeypatch.setattr(unit_refresh, '_load_helper', lambda path=None: ru)
return sudo
def test_web_side_lets_the_helper_decide_when_it_cannot_read_the_units(host, tmp_path, unreadable):
_stale(host)
result = _web(host, tmp_path, unreadable)
assert len(unreadable.calls) == 1
assert result['status'] == unit_refresh.REFRESHED and result['units'] == [ru.DISPLAY_UNIT]
def test_web_side_unreadable_and_already_current_is_current(host, tmp_path, unreadable):
result = _web(host, tmp_path, unreadable)
assert result['status'] == unit_refresh.CURRENT and result['message'] == ''
def test_web_side_unreadable_without_the_rule_asks_for_a_reinstall(host, tmp_path, unreadable, caplog):
unreadable.rc, unreadable.stderr = 1, 'sudo: a password is required'
result = _web(host, tmp_path, unreadable)
assert result['status'] == unit_refresh.NEEDS_REINSTALL
assert 'could not be checked' in result['message']
def test_web_side_trusts_the_helper_about_what_changed(host, tmp_path):
"""An older installed helper that renders differently changed nothing: nothing to roll back."""
_stale(host)
def older_helper(args, **kwargs):
return subprocess.CompletedProcess(args, 0, stdout='units: up to date\n', stderr='')
result = _web(host, tmp_path, older_helper)
assert result['status'] == unit_refresh.CURRENT
def test_web_side_without_the_helper_asks_for_a_reinstall(host, tmp_path, caplog):
_stale(host)
sudo = Sudo()
result = _web(host, tmp_path, sudo, helper_installed=False)
assert result['status'] == unit_refresh.NEEDS_REINSTALL and sudo.calls == []
assert 'first_time_install.sh' in result['message']
assert 'reinstall' in caplog.text
@pytest.mark.parametrize('stderr', [
'sudo: a password is required',
'Sorry, user ledpi is not allowed to run \'/usr/local/sbin/ledmatrix-refresh-units\' as root on ledpi.',
])
def test_web_side_without_the_sudo_rule_asks_for_a_reinstall(host, tmp_path, stderr, caplog):
_stale(host)
result = _web(host, tmp_path, Sudo(rc=1, stderr=stderr))
assert result['status'] == unit_refresh.NEEDS_REINSTALL
assert 'configure_web_sudo.sh' in result['message']
assert 'no sudo rule' in caplog.text
def test_web_side_reports_a_helper_refusal_as_a_failure(host, tmp_path):
_stale(host)
result = _web(host, tmp_path, Sudo(rc=1, stderr='ledmatrix-refresh-units: systemd/x refusing'))
assert result['status'] == unit_refresh.FAILED
assert 'refusing' in result['message']
def test_web_side_on_a_machine_without_the_units_does_nothing(tmp_path):
sudo = Sudo()
result = unit_refresh.refresh_after_update(run=sudo, systemd_dir=str(tmp_path))
assert result['status'] == unit_refresh.SKIPPED and sudo.calls == []
def test_web_side_never_raises_on_a_broken_template(host, tmp_path):
host.set_template(ru.WEB_UNIT, host.template(ru.WEB_UNIT).replace('User=__USER__', f'User={OTHER_USER}'))
sudo = Sudo()
result = _web(host, tmp_path, sudo)
assert result['status'] == unit_refresh.FAILED and sudo.calls == []
+113 -7
View File
@@ -18,6 +18,7 @@ from src.plugin_system.resource_monitor import (
ResourceLimits, ResourceLimits,
ResourceLimitExceeded, ResourceLimitExceeded,
PSUTIL_AVAILABLE, PSUTIL_AVAILABLE,
METRICS_SNAPSHOT_KEY,
) )
@@ -174,7 +175,7 @@ class TestMetricsPersistenceChurn:
with patch.object(rm.time, "monotonic", return_value=12.0): with patch.object(rm.time, "monotonic", return_value=12.0):
mon.monitor_call("p", lambda: None) mon.monitor_call("p", lambda: None)
writes = [c for c in cache.set.call_args_list writes = [c for c in cache.set.call_args_list
if "plugin_metrics:" in str(c)] if c.args and c.args[0] == rm.METRICS_SNAPSHOT_KEY]
assert writes, \ assert writes, \
"the first snapshot was dropped because the process was young" "the first snapshot was dropped because the process was young"
@@ -184,7 +185,7 @@ class TestMetricsPersistenceChurn:
for _ in range(50): for _ in range(50):
mon.monitor_call("p", lambda: None) mon.monitor_call("p", lambda: None)
writes = [c for c in cache.set.call_args_list writes = [c for c in cache.set.call_args_list
if c.args and str(c.args[0]).startswith("plugin_metrics:")] if c.args and c.args[0] == METRICS_SNAPSHOT_KEY]
assert len(writes) == 1, ( assert len(writes) == 1, (
f"50 calls produced {len(writes)} metric writes; expected 1") f"50 calls produced {len(writes)} metric writes; expected 1")
@@ -194,10 +195,10 @@ class TestMetricsPersistenceChurn:
mon = PluginResourceMonitor(cache, enable_monitoring=False) mon = PluginResourceMonitor(cache, enable_monitoring=False)
mon.monitor_call("p", lambda: None) mon.monitor_call("p", lambda: None)
# pretend the interval has passed # pretend the interval has passed
mon._metrics_persisted_at["p"] -= rm._METRICS_PERSIST_INTERVAL + 1 mon._snapshot_persisted_at -= rm._METRICS_PERSIST_INTERVAL + 1
mon.monitor_call("p", lambda: None) mon.monitor_call("p", lambda: None)
writes = [c for c in cache.set.call_args_list writes = [c for c in cache.set.call_args_list
if c.args and str(c.args[0]).startswith("plugin_metrics:")] if c.args and c.args[0] == METRICS_SNAPSHOT_KEY]
assert len(writes) == 2 assert len(writes) == 2
def test_in_memory_metrics_stay_exact_while_writes_are_skipped(self): def test_in_memory_metrics_stay_exact_while_writes_are_skipped(self):
@@ -213,8 +214,11 @@ class TestMetricsPersistenceChurn:
mon.reset_metrics("p") mon.reset_metrics("p")
mon.monitor_call("p", lambda: None) mon.monitor_call("p", lambda: None)
writes = [c for c in cache.set.call_args_list writes = [c for c in cache.set.call_args_list
if c.args and str(c.args[0]).startswith("plugin_metrics:")] if c.args and c.args[0] == METRICS_SNAPSHOT_KEY]
assert len(writes) == 2, "reset should clear the throttle timestamp" # The first call, the reset (which drops the plugin), the next call.
assert len(writes) == 3, "reset should clear the throttle timestamp"
assert "p" not in writes[1].args[1]["plugins"]
assert writes[2].args[1]["plugins"]["p"]["call_count"] == 1
def test_a_failed_write_does_not_buy_the_next_interval_of_silence(self): def test_a_failed_write_does_not_buy_the_next_interval_of_silence(self):
"""A set() that raises must not count as having persisted. """A set() that raises must not count as having persisted.
@@ -230,5 +234,107 @@ class TestMetricsPersistenceChurn:
# the very next call must try again rather than skip the interval # the very next call must try again rather than skip the interval
mon.monitor_call("p", lambda: None) mon.monitor_call("p", lambda: None)
writes = [c for c in cache.set.call_args_list writes = [c for c in cache.set.call_args_list
if c.args and str(c.args[0]).startswith("plugin_metrics:")] if c.args and c.args[0] == METRICS_SNAPSHOT_KEY]
assert len(writes) == 2, "a failed write should be retried, not skipped" assert len(writes) == 2, "a failed write should be retried, not skipped"
class TestOneSnapshotForAllPlugins:
"""Every plugin's metrics share one record, written at most once a minute.
A record per plugin, each throttled to 30 s, was still two writes a minute
per plugin. The web UI's output must not change: it reads the same numbers
for the same plugins, from the snapshot or, for a plugin the snapshot does
not have yet, from the per-plugin record an older version left.
"""
@pytest.fixture
def cache_dir(self, tmp_path):
return str(tmp_path)
def _manager(self, cache_dir):
from src.cache_manager import CacheManager
with patch('src.cache_manager.CacheManager._get_writable_cache_dir',
return_value=cache_dir):
manager = CacheManager()
manager.stop_cleanup_thread()
return manager
def test_many_plugins_one_write(self):
cache = _cache()
mon = PluginResourceMonitor(cache, enable_monitoring=False)
for _ in range(10):
for pid in ("a", "b", "c", "d"):
mon.monitor_call(pid, lambda: None)
sets = cache.set.call_args_list
assert len(sets) == 1
assert not any(str(c.args[0]).startswith("plugin_metrics:") for c in sets)
def test_the_web_reads_what_the_display_has(self, cache_dir):
display = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
for pid in ("a", "b"):
display.monitor_call(pid, lambda: None)
display._snapshot_persisted_at = None # let the next call publish
display.monitor_call("a", lambda: None)
web = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
for pid in ("a", "b"):
assert web.get_metrics_summary(pid, force_reload=True) == \
display.get_metrics_summary(pid)
assert web.get_metrics_summary("a", force_reload=True)["call_count"] == 2
def test_a_per_plugin_record_from_an_older_version_is_still_read(self, cache_dir):
old = self._manager(cache_dir)
old.set("plugin_metrics:legacy", {"call_count": 9, "total_execution_time": 1.8,
"last_update_time": time.time()})
display = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
display.monitor_call("other", lambda: None)
web = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
assert web.get_metrics_summary("legacy", force_reload=True)["call_count"] == 9
# The display carries the count on from it, into the snapshot.
display.monitor_call("legacy", lambda: None)
display._snapshot_persisted_at = None
display.monitor_call("other", lambda: None)
assert web.get_metrics_summary("legacy", force_reload=True)["call_count"] == 10
def test_a_restart_keeps_plugins_it_has_not_run(self, cache_dir):
first = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
first.monitor_call("disabled_later", lambda: None)
restarted = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
restarted.monitor_call("running", lambda: None)
web = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
assert web.get_metrics_summary("disabled_later", force_reload=True)["call_count"] == 1
assert web.get_metrics_summary("running", force_reload=True)["call_count"] == 1
def test_a_reset_from_the_web_sticks_for_a_plugin_the_display_is_not_running(
self, cache_dir):
display = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
display.monitor_call("idle", lambda: None)
web = PluginResourceMonitor(self._manager(cache_dir), enable_monitoring=False)
assert web.get_metrics_summary("idle", force_reload=True)["call_count"] == 1
web.reset_metrics("idle")
display._snapshot_persisted_at = None
display.monitor_call("busy", lambda: None)
assert web.get_metrics_summary("idle", force_reload=True)["call_count"] == 0
def test_a_long_idle_plugin_is_dropped(self, cache_dir):
import src.plugin_system.resource_monitor as rm
manager = self._manager(cache_dir)
manager.set(METRICS_SNAPSHOT_KEY, {"schema": 1, "plugins": {
"gone": {"call_count": 3, "last_update_time":
time.time() - rm._METRICS_SNAPSHOT_ENTRY_MAX_AGE - 10},
"recent": {"call_count": 4, "last_update_time": time.time() - 60},
}})
display = PluginResourceMonitor(manager, enable_monitoring=False)
display.monitor_call("p", lambda: None)
plugins = manager.get(METRICS_SNAPSHOT_KEY, max_age=None, memory_ttl=0)["plugins"]
assert set(plugins) == {"recent", "p"}
@pytest.mark.parametrize("junk", [[1, 2], {"schema": 99, "plugins": {"p": {}}},
{"schema": 1, "plugins": "nope"}])
def test_an_unusable_snapshot_is_ignored(self, junk):
cache = MagicMock()
cache.get.side_effect = lambda key, **kw: junk if key == METRICS_SNAPSHOT_KEY else None
mon = PluginResourceMonitor(cache, enable_monitoring=False)
assert mon.get_metrics_summary("p", force_reload=True)["call_count"] == 0
mon.monitor_call("p", lambda: None)
written = cache.set.call_args.args[1]
assert written["plugins"]["p"]["call_count"] == 1

Some files were not shown because too many files have changed in this diff Show More