Commit Graph
2078 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5.5 01b82edc23 fix(display): capture-mode check tolerates a DisplayManager built without __init__
main's new test_display_manager_logging builds one with object.__new__;
set_scrolling_state here asks whether the thread is capturing, which read
the per-thread state __init__ creates. Read it the way _current_surface()
already does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:49:07 -04:00
Chuck 9c81e6b42f Merge branch 'claude/hdpi-scroll-performance-antialiasing-4ae609' into claude/offscreen-rendering
# Conflicts:
#	src/display_manager.py
#	src/vegas_mode/config.py
#	src/vegas_mode/render_pipeline.py
2026-09-24 17:46:58 -04:00
ChuckandClaude Opus 5.5 c43c10bd78 fix(vegas): keep the sub-pixel frame interval #637 removed as dead code
On main, RenderPipeline._frame_interval was set and never read, so #637
dropped it. Here the frame_interval property reads it for the sub-pixel
path (the crisp path solves its own), so the merge brings it back, with a
comment saying who reads it. #637's new coordinator test drives a
MagicMock pipeline, which this branch's pacing reads frame_interval and
target_fps from; give the mock real ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:42:42 -04:00
Chuck 8482af7814 Merge commit 'b9416ef803028ec37e0f2f756c051e52d52622d7' into claude/hdpi-scroll-performance-antialiasing-4ae609
# Conflicts:
#	src/vegas_mode/config.py
#	src/vegas_mode/coordinator.py
2026-09-24 17:38:08 -04:00
ChuckandClaude Opus 5.5 b9416ef803 fix(display): Vegas teardown and 240s default; remove dead Vegas buffer code (#637)
* fix(display): tear down Vegas mode on controller cleanup

DisplayController.cleanup() never called VegasModeCoordinator.cleanup(),
so the Vegas teardown (stop, pipeline/stream reset, adapter cache drop)
was unreachable. Call it before the display manager is cleaned up, and
skip it when Vegas was never created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): default max_cycle_duration to the documented 240s

The template, the web UI help, CONFIG_REFERENCE and the controller all
say 240, but the code defaulted to 600 in two places, so a config
without the key ran Vegas iterations 2.5x longer than documented.

from_config now falls back to the dataclass field defaults instead of
repeating each one, so the two copies can no longer drift, and the
controller's follower scroll-speed default reads VegasModeConfig's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): let run.py -d show display_manager's DEBUG output

display_manager pinned its logger to INFO at import, overriding the root
level, so debug mode never showed its DEBUG lines. Use get_logger() from
src.logging_config like the rest of the core and leave the level to the
logging setup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): run each startup validation check once

StartupValidator.validate_all() ran twice at boot, before and after the
plugin manager was created, so every config, cache, display and
systemd-unit warning was logged twice. The second pass now runs only the
plugin checks. Drop the commented-out raise_on_errors line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): one INFO line per plugin-list refresh

StreamManager logged "=" * 60 banners and a line per plugin (INCLUDED,
SKIPPED, FETCHING CONTENT, SEGMENT CREATED) at INFO on every refresh and
fetch, i.e. at each cycle start and every 30s. Log one INFO summary of
the rotation per refresh and move the per-plugin detail, the weighting
breakdown and "no content this cycle" to DEBUG (the adapter still warns
when every content path fails).

Also drop the check/cross marks from the controller's log messages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): drop the per-iteration static-mode plugin scan

run_iteration() rebuilt _static_mode_plugins on every iteration, asking
every plugin for its display mode and logging the set at INFO, but
nothing ever read it: static pauses are triggered by
_check_static_plugin_trigger() from the next segment. Delete it, the
coordinator's get_ordered_plugins() that only it used, and the
write-only _static_pause_plugin / _static_pause_start.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): remove the staging buffer that was never filled

StreamManager and RenderPipeline carried a double-buffer design that
nothing used: _staging_buffer was only ever cleared or swapped, so
swap_buffers() never did anything and should_recompose()'s
staging_count > 0 branch was dead, and _active_scroll_image,
_staging_scroll_image, _is_rendering, _last_frame_time and
_frame_interval were written but never read. Delete the machinery and
rewrite the docstrings around what actually carries updates:
_pending_updates, consumed by process_updates() in swap mode and
invalidate_pending_updates() in continuous mode.

should_recompose() no longer builds a buffer-status dict every frame.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(display): tidy the display controller without changing behaviour

- Import VegasModeCoordinator locally instead of through module globals
  (there is no circular import to avoid).
- Drop hasattr() checks on attributes PluginManager.__init__ always sets
  (plugin_executor, plugin_last_update, get_plugin_lock,
  run_scheduled_updates*, stop_update_worker) and the dead "older
  manager" fallbacks; keep the health_tracker None checks, now via
  _health_tracker().
- Extract _display_once() for the per-frame display call both render
  loops copied, _advance_on_demand() for the two on-demand rotations,
  _reset_on_demand_fields() for the error and clear paths, and
  _timezone() / _in_window() for the two schedule checks.
- Remove always-true conditions and the unreachable non-plugin else
  branch in run(), and read _was_display_active / _last_published_mode /
  vegas_coordinator directly now that __init__ declares them.
- Declare the follower render state in __init__, name its tuning
  constants, add _follower_sign(), and share the 90/s sync send
  interval with the render pipeline (SYNC_SEND_INTERVAL).
- Delete history narration and the "Opt #N" labels, fix the comment
  that called _scroll_speed constant (hot reload updates it), and drop
  a startup timing log that measured nothing.
- render_pipeline / plugin_adapter: read display_manager.width/height
  as the properties they are, drop an empty TYPE_CHECKING block, an
  aliased threading import and a duplicated `if result and
  self.sync_manager:`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(display): trim dead code from display_manager

- Add _new_canvas() for the image/draw/fontmode="1" setup that was
  copied six times.
- Call resolve_double_sided() and compose_pixel_mapper_config() directly
  instead of through a module alias and a passthrough method, and replace
  the comment that said the passthrough read class attributes.
- Delete the unused _initialized flag and _ORIENTATION_ROTATE_DEGREES
  alias (no core or monorepo reader; tests stop resetting the flag), the
  test pattern's unreachable no-matrix branch (it only runs once the
  matrix exists), `del old_image  # help GC` (a no-op on a local), a
  duplicated early return in process_deferred_updates, and stale
  comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(vegas): remove unread fields and test-only helpers, fix docstrings

- ContentSegment: drop total_width, fetched_at, is_stale, image_count and
  is_static, none of which is read.
- StreamManager: drop _current_index (never advanced) and the test-only
  get_all_content_for_composition() and has_pending_updates();
  VegasModeConfig: drop the test-only is_plugin_included().
- geometry.find_blank_cut() has had no production caller since the crop
  moved to item boundaries; delete it and its tests.
- PluginAdapter: the _finalize docstring described separator_width
  between every image, and _crop_to_budget's said cuts snap to the
  nearest blank column; both now describe what the code does.
- Coordinator: the static-pause interrupt log no longer blames follower
  mode for every interrupt, and set_update_callback names the callback
  the controller actually wires.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(scroll): correct ScrollHelper comments and drop dead branches

- Four comments said the strip always starts with display_width of
  blank; it does only when lead_gap is None (Vegas passes its own).
- Delete the "Width calculation mismatch" warning: the image is created
  at the calculated width, so the two can never differ.
- Remove the two scroll_delay <= 0 fallbacks (which disagreed with each
  other): set_scroll_delay clamps it to at least 0.001 and nothing in
  core or the plugin monorepo assigns it directly.
- Trim the scipy history from the blend docstring.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(run): drop a redundant comment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): log set_scrolling_state only when it changes

Vegas and scrolling plugins set the scrolling state every frame, so once
display_manager's DEBUG output became visible in debug mode it printed
"Scrolling state set to: True" about 120 times a second. Log only when
the value differs from the previous one; the state, activity timestamp
and frame hold still update on every call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): display-vegas

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:37:15 -04:00
Chuck 478516dbd1 Merge branch 'claude/frame-timing-harness' into claude/offscreen-rendering 2026-09-24 17:37:07 -04:00
Chuck 690407860e Merge branch 'claude/hdpi-scroll-performance-antialiasing-4ae609' into claude/offscreen-rendering
# Conflicts:
#	CHANGELOG.md
2026-09-24 17:37:07 -04:00
ChuckandClaude Opus 5.5 f6afbdbb15 fix(web): Cache/Logs error mix-up, store errors, tab fallbacks; remove ~2.5k lines of dead JS (#639)
* fix(web): keep Cache and Logs helpers out of each other's way

Both partials declared top-level showError and escapeHtml. Their scripts
run at global scope after every HTMX swap, so whichever tab was opened
last owned window.showError, and a Cache failure after visiting Logs
rendered into the Logs panel (and the other way round). Each script is
now an IIFE; Cache still exports deleteCacheFile for its row buttons.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): make the HTMX-failure fallbacks for tab panels actually run

- The "HTMX never loaded" fallback read appElement.__x.$data, which is
  Alpine 2. The page ships Alpine 3, so the check was always false and
  the Overview never loaded without HTMX. It now reads Alpine.$data().
- The Overview and WiFi panels used hx-on::htmx:response-error, which
  htmx expands to "htmx:htmx:response-error", an event that never fires.
- loadTabContent sent requests with <body> as the source, so htmx fired
  its events on <body> and no panel's hx-on handler ran at all. The
  panel is now the source. htmx also resolves its promise on a 4xx/5xx,
  and the panel was stamped data-loaded anyway, leaving a skeleton that
  never retried; it is now stamped only when no responseError fired.

loadPluginsDirect, loadOverviewDirect and loadWifiDirect are merged into
one window.loadPartialDirect(id, url), which also runs the partial's
inline scripts before Alpine sees the markup (as htmx-config.js does on
htmx:afterSwap). The ~10 s "htmx never arrived" path in loadTabContent
uses it for every tab instead of four hard-coded ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): store and registry failures no longer wipe the Plugin Manager

showError replaced the whole #plugins-content with an error message, so
one failed store search, custom-registry install or saved-repository
call took the installed list, the store and every control with it, with
no way back short of reloading the tab. Those failures are now error
notifications. The full-panel message is kept only for a first load of
the installed list that failed (nothing to show yet); a failed refresh
of an already-rendered list is a notification too. showSuccess's
fallback branch, which wrote the message into innerHTML unescaped, is
gone: showNotification always exists.

The "Please try refreshing your browser" hint tested for the text
"Failed to Fetch", which no browser produces (Chrome says "Failed to
fetch", Firefox "NetworkError..."), so it never appeared. It now keys on
the failure itself: a TypeError from fetch(), or PluginAPI's
NETWORK_ERROR wrapper around one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): escape plugin action and install output on every path

executePluginAction escaped data.message and data.output when an action
failed but put data.message straight into innerHTML when it succeeded,
and set the OAuth step-2 button's innerHTML from the manifest's
step2_button_text. Plugin actions run plugin code, so that is plugin- or
server-controlled markup in the page. Both paths now escape, and the
button label is set with textContent.

The same pattern sat in the install-from-GitHub-URL status lines
(plugin_id, the server's message, and error.message, which can echo a
repository URL) and the custom-registry load error; those are escaped
too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): file-upload widget owns the image list and schedule editor

plugins_manager.js loads after the widget bundle, so its older copies of
deleteUploadedFile, updateImageList, hideUploadProgress, formatDate,
openImageSchedule, toggleImageScheduleEnabled, updateImageSchedule{Mode,
Time,Day} and updateCheckboxGroupData replaced the widget's. They are
deleted; the widget files are the only definitions.

Before switching over, the two sets were diffed and fixed so nothing
regresses:

- The old copy labelled the schedule/delete buttons for screen readers
  and lazy-loaded thumbnails; the widget now does both.
- The schedule button did nothing on a card rendered by
  plugin_config.html whenever the image id is a UUID (every upload): the
  template turns "-" into "_" in the editor's id, and neither JS copy
  did. Both now use the template's rule.
- The widget's "keep the open editor open" copied the editor's innerHTML
  into the new list. That dropped its event listeners and showed the old
  values, so after the first change the editor looked live but ignored
  input. A schedule edit now saves to the hidden input and updates the
  card's summary in place without re-rendering the list; a list re-render
  (upload, delete) rebuilds an open editor from the data. Editor controls
  are routed by one delegated change listener, so there are no
  per-element listeners to lose.
- The old deleteUploadedFile had a JSON branch that removed a
  #file_<id> element and skipped the re-render. No template or script
  renders such an element, and JSON uploads are listed through
  updateImageList like images, so re-rendering (the widget's behaviour) is
  the consistent one; the branch was not carried over.
- The template always renders the summary line (".image-schedule-summary",
  "Always shown" when unscheduled) so an edit has a line to update.

The inline-handler test evaluated plugins_manager.js's updateImageList;
test_file_upload_widget.js now covers the widget's list and editor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete the unused handleCredentialsUpload

Its last caller went when plugin_config.html switched credential uploads
to the file-upload widget's handleSingleFileSelect. Nothing in the web
UI, the tests or the plugin monorepo references it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete dead and shadowed front-end code

Nothing calls any of these (checked across web_interface/, test/ and the
ledmatrix-plugins monorepo, including hx-*/x-*/onclick attributes):

- app-shell.js: the Alpine methods refreshPlugins (it called a
  nonexistent this.searchPluginStore), loadPluginConfig,
  savePluginConfig, getSchemaPropertyType, escapeCssSelector,
  formatCommitInfo and formatDateInfo, and the top-level copies of
  savePluginConfig, getSchemaPropertyType, escapeCssSelector,
  formatCommitInfo, formatDateInfo and togglePluginFromTab. Plugin config
  forms save through hx-post in plugin_config.html.
- window.reconnectSSE (app-shell.js); window.updateArrayTableAddButtonState
  (array-table.js).
- toggleNestedSection, defined twice (app-shell.js and
  plugins_manager.js) and called from nowhere.
- plugins_manager.js: the window.initializePlugins wrapper around an
  IIFE-local origInit that was always undefined, and __pluginDomReady,
  which was written but never read.
- display.html's fixInvalidNumberInputs fallback: app-shell.js defines it
  before any partial loads.
- base.html's window.loadCodeMirror and the two CodeMirror stylesheet
  preloads, and the .CodeMirror rules in plugins.html. The raw JSON
  editor is a plain textarea.

Also deleted: app-shell.js definitions that a later script always
replaced, so they never ran: executePluginAction (plugins_manager.js
assigns its own), uninstallPlugin and its pollUninstallOperation
(plugins_manager.js), and updateAllPlugins (install_manager.js).

vendor/codemirror stays: test/test_web_smoke.py still requests
codemirror.min.js as a sample static asset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): call showNotification without checking it exists

app-shell.js defines window.showNotification (a stand-in that queues
until the notification widget loads) and base.html runs it, deferred,
before every other script that notifies: app.js, the utilities, the
widget bundle, plugins_manager.js, and all partials, which HTMX loads
after the page. The 81 `typeof showNotification === 'function'` /
`!== 'undefined'` checks, the `window.showNotification || console.log`
and `|| alert` fallbacks, and their else branches (alert(), console
output, and schedule.html's own hand-built toast) could never take the
fallback path. They are removed, as is fonts.html's second copy of the
queueing stand-in.

The stand-in in app-shell.js keeps its guard (it must not replace the
widget's implementation if load order ever changes), and BaseWidget's
public notify()/getNotificationFunction() keep their shape for widgets
that plugins ship.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one HTML escaper, window.LEDEscape

About 30 files each carried their own escapeHtml / escapeAttr / escHtml /
_esc / escapeJs. They disagreed: several (notification.js, display.html's
escapeHtml, operation_history.html, the app() stub) did not escape
quotes, google-calendar-picker.js and tools.html's escHtml left ' alone,
and some turned 0 into ''. Most were fine only because the quote-safe
widget copies were preferred at runtime.

window.LEDEscape now lives at the top of app-early.js, a blocking script
in <head>, so it exists before any other script runs:

  html(v)          & < > " ' as entities, null/undefined as ''
  attr(v)          the same, for call sites that want to say "attribute"
  jsStringAttr(v)  a JS string literal safe inside an inline handler

Every former copy is now a one-line name for it (kept so call sites do
not change), widgets included, with no fallback. plugins_manager.js
loses its four escapeJs wrappers (callers use jsStringAttr), the
duplicate escapeAttr and escapeHtml inside renderInstalledCards and
renderCustomRegistryPlugins, and the window.escapeHtml /
window.escapeAttribute exports, which nothing read.
addArrayObjectItem's fallback markup (with a sixth hand-written escape
chain) is gone too: window.renderArrayObjectItem is defined earlier in
the same file, so the fallback could not run. The unused escapeHtml
methods on the Alpine app (app-early.js stub and app-shell.js) are
deleted.

test_html_escaping.js now runs LEDEscape and every remaining name for it,
and fails if a hand-rolled escaper reappears anywhere in web_interface/.
Suites that evaluate slices of plugins_manager.js or widget files load
LEDEscape from app-early.js through test/js/led_escape.js.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): stop htmx re-running partial scripts after every tab load

htmx-config.js runs each swapped-in <script> itself on htmx:afterSwap
and meant to turn htmx's own script handling off with
htmx.config.allowScriptTags = false. It did that once, while setting up,
but base.html loads htmx with a dynamic <script>, so htmx was not defined
yet and the setting never applied. On every tab load htmx then tried to
run each script again in its settle phase, found it already replaced
(no parent node) and threw "Cannot read properties of null (reading
'insertBefore')" into the console, which also skipped the rest of that
swap's settle tasks.

The setting is now applied in the afterSwap handler, which always runs
after htmx exists and before htmx settles the same swap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): show "--" for a system stat the server could not read

The stats stream and /system/status now send null for a metric they
cannot read (cpu_temp off a Pi, for one) instead of 0. updateSystemStats
built the header and Overview text as value + unit, so a null showed as
"null°C". CPU, memory and temperature, in the header and on the
Overview, now render "--" plus the unit for null or a missing field --
the same placeholder the page starts with, and what tools.html already
shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one Alpine accessor and one plugin-list signal

window.getApp() (app-early.js) returns the root <body x-data="app()">
component through Alpine's public Alpine.$data, or null before Alpine
has initialised it. It replaces the private el._x_dataStack[0] reads in
app.js, app-early.js, app-shell.js, settings-search.js, overview.html and
plugins_manager.js, the three local getAppComponent/appData/getAppData
copies, and the Alpine 2 el.__x.$data fallbacks, which Alpine 3 never
provides.

Publishing the installed-plugin list: one load set window.installedPlugins
and dispatched pluginsUpdated twice (loadInstalledPlugins, then
renderInstalledPlugins), then wrote into the Alpine component through
_x_dataStack[0] and called its updatePluginTabs() directly, and
app-early.js's global listener set window.installedPlugins a third time
and called updatePluginTabs() again. Now renderInstalledPlugins is the
one publisher: it sets window.installedPlugins and dispatches
pluginsUpdated once, and the full app()'s listener (app-shell.js) is the
receiver. The app-early.js listener only builds the tab row while the app
is not the full implementation yet. The "grid not loaded yet" case is a
normal state (Plugin Manager tab not opened), so it logs through
pluginLog instead of console.warn.

updatePluginTabs had a "Debounce" comment and clearTimeout over a timer
nothing ever set, and two identical branches; it now just calls
_doUpdatePluginTabs (app-early.js detects the full implementation by
that name in its source, which the new comment says).

app()'s unused baseComponent lookup is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): reload the plugin list after installs and failed toggles

Several callers refreshed the installed list with

    if (typeof loadInstalledPlugins === 'function') loadInstalledPlugins();
    else if (typeof window.loadInstalledPlugins === 'function') ...

but loadInstalledPlugins is local to the plugin-manager IIFE and
window.loadInstalledPlugins is never defined, so from outside that IIFE
both tests were false and nothing reloaded:

- A failed plugin toggle left the switch drawn in the new state while
  the data said the old one. It now re-renders from the reverted data.
  The optimistic in-place edit also has to forget the grid's
  last-rendered markup, or setGridHtmlIfChanged sees identical HTML and
  skips the revert. A successful toggle still keeps the switch (and
  focus) as drawn.
- Installing from a GitHub URL (the early handleGitHubPluginInstall),
  installing or uploading a Starlark app, and toggling a Starlark app on
  its config tab never refreshed the list, so the new app had no tab or
  Installed badge until the page was reloaded. They now force a reload
  through window.pluginManager.loadInstalledPlugins(true), and the
  Starlark grid redraws when that finishes instead of after a fixed
  500 ms.
- The Starlark uninstall inside the IIFE reloaded from the 3 s cache,
  which could still hold the app; it now forces a reload.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): route debug output through debugLog

base.html defines window.debugLog, gated on localStorage.pluginDebug.
plugins_manager.js read the same key twice more into its own flags
(_PLUGIN_DEBUG_EARLY, and PLUGIN_DEBUG behind a pluginLog() wrapper), and
api_client.js's RequestThrottler had a separate `debug` property with a
setDebug() that nothing called. All of it now goes through debugLog. The
"functions defined" dumps with their ✓ lines, and two per-plugin
"enabled=" loops that ran on every render, are dropped; "[PLUGINS STUB]"
labels on code that has not been a stub for a long time read
"[PLUGINS]".

Ungated console.log calls that announced normal events on every page
load or action (settings search and tooltips registering, every toast
repeated to the console, the schedule pickers initialising, widget
registry unregister/clear) go through debugLog too. What remains on
console.log is the widget registry's on-demand LEDMatrixWidgets.debug()
dump and BaseWidget.notify's no-notifier fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop waits and guards that could never fire

- handlePluginAction polled up to 10 x 50 ms for window.togglePlugin,
  configurePlugin, updatePlugin and uninstallPlugin before calling them.
  All four are defined when the scripts load, before any card can be
  clicked, so the poll always succeeded at once; it now calls them.
  The long thinking-aloud comment over the toggle state is replaced by
  two lines on why the stored state, not the checkbox, decides.
- initializePlugins checked typeof on setupGitHubInstallHandlers and
  applyStoreFiltersAndSort, function declarations in the same IIFE, and
  wrapped window.checkGitHubAuthStatus(), which returns a promise with
  its own .catch, in try/catch.
- searchPluginStore wrapped each "#store-count" update (a getElementById
  and an innerHTML assignment) in try/catch four times; one
  setStoreCount() helper does it. The store's post-render re-attach of
  the GitHub token handler dropped its try/catch and existence checks
  for the same reason.
- The load-time fallback outside the IIFE tested typeof
  initializePluginPageWhenReady, which is IIFE-local and so always
  undefined there; it calls window.initPluginsPage directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete two unused plugin-manager helpers

stopOnDemand (IIFE-local; the page's stop button calls window.stopOnDemand
from app-shell.js) and debounce had no callers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): document the plugin-config handlers templates call

validatePluginConfigForm, handleConfigSave, handleToggleResponse,
handlePluginUpdate and refreshPluginConfig each get a JSDoc naming the
attribute in partials/plugin_config.html that calls it and what the
return value means (only validatePluginConfigForm's matters: false
cancels the submit).

- The `if (!window.__pluginConfigHandlersInitialized)` wrapper is gone:
  app-shell.js runs once per page, so it was never false. The block is
  dedented one level; `git diff -w` shows the real change.
- The three handlers read xhr.responseJSON first. XMLHttpRequest has no
  such property (it is jQuery's), so that branch never ran; one
  xhrJson(xhr) helper parses responseText for all of them, with the same
  fallbacks as before.
- runPluginOnDemand and stopOnDemand checked that plugins_manager.js's
  openOnDemandModal/requestOnDemandStop exist; plugins_manager.js is on
  every page, so they call them.
- fixInvalidNumberInputs had a stray "Notification helper function"
  comment on top of its own; a leftover "section toggle ... duplicate
  definition removed" note is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): one toast per save, and a failed durations save says so

app.js's global htmx:afterRequest listener showed the server's message
for every htmx request, and every form and button that posts through
htmx (plugin config save/toggle/update, Display, Durations, General,
Schedule, Dim schedule, the Overview actions) also reports its own result
from hx-on after-request. Each save showed two toasts. The global
listener now stays quiet for a request whose element, or its form, has
its own after-request handler.

That exposed the Rotation & Durations form's handler, which read
xhr.responseJSON: XMLHttpRequest has no such property, so it always said
"Durations saved" in green, even when the save failed (the global toast
had been the only place the error showed). display.html already had a
correct version (2xx only counts as saved; the server's message wins;
its status may refine success but never overturn failure). That is now
window.showSaveResult(xhr, savedText, failedText) in app.js, used by the
Display, Durations and General forms; General's inline copy of the same
logic is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(web): file headers and comments that say what the code does now

- plugins_manager.js, app-shell.js, app.js and app-early.js open with a
  header: what the file owns, how base.html loads it and in what order
  relative to the others, and the globals it defines. app-early.js's
  app() stub also says why it exists and that, with app-shell.js now
  loaded before Alpine, it does not run in practice.
- base.html's note on plugins_manager.js said it must load last to win
  over same-named functions in app.js/app-shell.js; there are none left,
  so it now gives the real reason (it uses everything loaded before it).
- Change-narration and "already defined at the top, no need to redefine"
  notes are gone or rewritten as present-tense reasons; comments that
  were wrong are fixed ("Toggle password visibility" over the function
  that opens the token panel, "Insert before the closing </nav>" over an
  appendChild, "(from v2)", the export note that still listed
  escapeHtml). About forty comments that restated the line below them
  are removed, and a second window.currentPluginConfig = null outside the
  IIFE is dropped (the IIFE sets it).
- The file-upload, checkbox-group and custom-feeds widgets' render()
  stubs say plainly that the widget is rendered server-side, instead of
  "for now" / "placeholder for future client-side rendering".

test_plugin_action_delegation.js sliced the source up to one of the
removed notes; it now ends the slice at the next section header.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): keep the escapeHtml/escapeAttribute globals for plugin pages

6da77363 removed window.escapeHtml and window.escapeAttribute because nothing in core or the plugin monorepo read them. Plugin web UIs served through serve_plugin_web_ui and third-party plugin pages may still call them, so they come back as aliases of window.LEDEscape.html and .attr, defined in app-early.js before any other script runs. test_html_escaping.js checks the aliases exist.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web-frontend

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): encode the image thumbnail path; match script tags case-insensitively

CodeQL flagged the upload widget building an <img> src from a stored path,
and the escaper test extracting inline scripts with a case-sensitive regex.
Each path segment is now URL-encoded (still a same-origin path, and correct
for names with spaces or

* fix(web): clear Codacy findings in the escaper, app shell and upload widget

- LEDEscape looks entities up in a Map instead of indexing an object.
- showNotification is declared as a global for app-shell.js.
- openImageSchedule checks the index is a non-negative integer and reads
  the image with Array.prototype.at.
- The schedule editor calls escapeHtml directly and documents why its
  innerHTML template is safe: every value is escaped or constrained.
  The remaining rule hits are suppressed on that line with the reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): build the image schedule editor with DOM calls

Codacy does not honour inline suppressions, and the editor's innerHTML
template kept tripping its XSS rules even though every value was escaped.
The editor is now built with a small element helper (createElement and
setAttribute), so no value is ever parsed as HTML, and the file's own
escapeHtml goes away.

Also for Codacy:
- LEDEscape.attr is its own function rather than a second name for html.
- The tab loader records a failed load on the panel (data-load-failed)
  from a named handler, instead of a closure over a local flag.

The fake DOM in test_file_upload_widget.js gains append/replaceChildren,
its hostile-id check now asserts the id arrives as attribute data with no
innerHTML anywhere in the editor, and test_html_escaping.js drops the
file-upload.js escaper it no longer has.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): schedule editor helpers as plain functions

Codacy's lint flags arrow functions held in local constants and a forEach
callback that returns a value. The editor's pieces are now named function
declarations (displayStyle, scheduleModeOption, scheduleRangeTime,
scheduleDayTime, scheduleDayRow) taking what they need as arguments, and
the element helper loops with for...of. htmx is declared as a global in
app-shell.js. Output is unchanged; test_file_upload_widget.js passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:36:53 -04:00
Chuck 99a3608516 Merge remote-tracking branch 'origin/main' into claude/frame-timing-harness 2026-09-24 17:36:48 -04:00
Chuck 42ff3825b4 Merge remote-tracking branch 'origin/main' into claude/hdpi-scroll-performance-antialiasing-4ae609 2026-09-24 17:36:47 -04:00
Chuck edfcd9e2a1 Merge remote-tracking branch 'origin/main' into claude/frame-timing-harness
# Conflicts:
#	CHANGELOG.md
2026-09-24 17:35:58 -04:00
ChuckandClaude Opus 5.5 3a81f38f09 fix(web): uniqueItems saves, /health count, Vegas order wipe; one list-repair helper (#638)
* fix(web): drop repeats from uniqueItems lists before validating a plugin save

dedup_unique_arrays lost its only caller in #330, so submitting a value a
uniqueItems list already holds (a stock symbol saved once and posted again)
failed the whole save with a validation error. _prepare_plugin_config_for_save
runs it again just before validation, which covers both POST /plugins/config
and plugin sections posted to /config/main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): /health counts the discovered plugins and logs the checks it fails

The plugin check counted plugin_manager.get_available_plugins(), which
PluginManager does not have, behind a hasattr guard that made plugin_count 0
on every device. It now counts the discovered manifests, discovering first
when nothing has been scanned yet.

The config, plugin and hardware checks answered "see logs for details"
without logging anything. Each now logs a warning with the traceback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): store refresh no longer claims a commit-metadata refresh

POST /plugins/store/refresh read fetch_commit_info (or fetch_latest_versions)
only to append "(with refreshed commit metadata from GitHub)" to its message.
It never fetched any: the route re-downloads the registry and nothing else.
search_plugins takes the flag, but it reads commit info through its cache,
so passing it on would not refresh anything either. The flag is ignored now
and the message says what happened.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): refuse a malformed Vegas plugin order instead of clearing it

A vegas_plugin_order or vegas_excluded_plugins value that was not JSON, or
not a list, was stored as [] and the save answered 200, so a bad value wiped
the saved order or exclusions. Both now answer 400 and save nothing, the way
plugin_rotation_order already did; the three share one parser. A list that
holds anything but plugin-id strings is refused as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): per-plugin health and metrics read the display service's latest

GET /plugins/health/<id> and /plugins/metrics/<id> called get_health_summary
and get_metrics_summary without force_reload, so they answered with whatever
the web process read first and kept in memory, while the display service kept
writing newer state. They now pass force_reload=True, as the list routes do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin config reset saves through the shared atomic save

POST /plugins/config/reset called config_manager.save_config directly, so it
took no backup, and a failed write escaped as an unhandled exception. It then
handed on_config_change the raw stored section, not the prepared config a
loaded plugin runs with. It now saves through _save_config_atomic with a
backup, answers CONFIG_SAVE_FAILED when that fails, and notifies with
_prepared_plugin_config, as POST /plugins/config does.

POST /plugins/toggle carried its own copy of _save_config_atomic's
save_config_atomic-or-save_config fallback; it calls the shared helper now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): one reading and one "unavailable" for each system metric

system_metrics.collect_system_metrics() promised None for a metric it could
not read, but returned cpu_temp as 0 off a Pi, and the whole no-psutil
fallback as zeros. GET /system/status measured the same numbers a second time
with its own code, and answered None there. Now both come from
collect_system_metrics(), and "unavailable" is None everywhere.

/system/status keeps its 0.1s CPU sample and its 10s cache, and gains
nothing it did not already send. Two differences: without psutil it answers
200 with null metrics instead of 503, and a disk it cannot stat is null
instead of a 500.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): /display/current sends the snapshot as-is and logs a failed read

GET /display/current PIL-decoded the preview snapshot and re-encoded it before
base64-ing it, spending CPU on the Pi to send the same picture, and dropped
any failure with `except Exception: pass`. The /stream/display SSE stream
already passed the PNG's bytes straight through.

Both now read through web_interface/display_preview.py and answer with the
same payload. A missing snapshot is still a null image; any other read
failure is logged as a warning. /health reads the snapshot path from the same
module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one helper puts a submitted plugin config's lists back

The plugin-config save turned position-keyed dicts ({"0": ..., "1": ...})
back into lists in five copies: four in the form path's
fix_array_structures (whose prefix branches never ran, since no caller
passed one), and _fix_json_arrays on the JSON path. It then force-fixed
the news plugin's feeds.custom_feeds by name, in case the generic pass had
missed it. src/web_interface/config_arrays.coerce_array_shapes now does it
for both paths, custom_feeds included. ensure_array_defaults duplicated
_fix_none_arrays and is gone.

In the same function: the union-type re-checks that the null handling
above them made unreachable, the "(temporary)" random_seed debug log, and
a commented-out log line are removed. A failed validation is logged once
as a warning, not four ERROR lines and a WARNING.

Element types are left to normalize_config_values, which already converted
them for both paths. One difference: the form path no longer adds an empty
{} for a nested object the post left out that has no defaults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): import at module top and log through the module logger

The web_interface.cache imports in config.py and fonts.py were wrapped in
`except ImportError` fallbacks. It is an in-repo module that imports nothing
from the project, so it cannot fail to import; it is imported once at module
top, as system.py now does. cache.py's docstring said blueprints import it
lazily "to avoid circular imports"; it now says why that is unnecessary.

Five logging.error calls in the dim-schedule GET and three logging.warning
calls in plugins.py went to the root logger; they use the module logger.
Function-local re-imports of json, os, shutil, logging and Path, all
already imported by the module, are gone. The `import os` inside two except
blocks of save_plugin_config also made os a local name for the whole function.

execute_plugin_action's step-1 handler gets a comment saying why it stays:
it looks like a copy of the blueprint handler, but without it a
TimeoutExpired from the plugin's script would reach the route's own
`except subprocess.TimeoutExpired` and be answered as a 408.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): app.py loses dead CSRF and reconciliation state, comments fixed

- csrf was always None, so `if csrf: csrf.exempt(...)` never ran, and its
  note that the api_v3 blueprint "is exempted above" named an exemption that
  does not exist. Both are gone; the reason there is no CSRF protection stays,
  shortened.
- The SSE rate-limit comment called the default "tight" at 20 per minute. The
  default is 1000 per minute and the streams' 200 is the tighter one; the
  comment now says so. The limits are unchanged.
- _reconciliation_done was written and never read. The docstring that
  explains why reconciliation runs once keeps its reason, in the present
  tense.
- Removed: a dangling "import cache functions" comment with no import under
  it, a "security check ... within project_root" label on an existence check,
  the "(simplified version)" narration, and the note that no redirect route is
  needed. The preview loop's sleep comment no longer mentions a PIL encode
  that the loop does not do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(web): api_v3 comments name the package __init__, not a _common module

Every route module's docstring said the shared blueprint comes "from
._common", a module the package split never created; they name the
package __init__. The PROJECT_ROOT comment described the path from
_common.py; it now describes this package and keeps the incident it
guards against. The "(corrected) in this commit" note in
resolve_pull_command and the /health comment the split's mechanical
time -> _pkg.time rewrite garbled ("Stamp the start _pkg.time") read
correctly again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop hasattr checks for attributes PluginManager always has

PluginManager.__init__ sets health_tracker and resource_monitor (to None
until they are configured), so the seven
hasattr(api_v3.plugin_manager, ...) guards in the health, metrics and limits
routes were always true. The falsy checks that do the work stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): pages_v3 dispatches partials from a dict with one error handler

load_partial chose a loader through a fourteen-branch if/elif, and thirteen
of the loaders then wrapped themselves in the same try/except, logging
"Error loading partial" without saying which. The route now looks the name up
in _PARTIAL_LOADERS and has the one handler, which logs the partial's name.
The loaders just render. _load_tools_partial keeps its own messages. The
search index's _partial_html already catches a loader that raises.

serve_plugin_web_ui repeated _plugin_dir_for inline (containment plus the
ledmatrix- prefix fallback); it calls it now. Also removed: the unused
markupsafe.escape import, function-local json/Path re-imports, and unused
exception bindings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): remove unused imports, locals and a try that cannot fail

- get_error_aggregator was imported by the api_v3 package and used by no
  one; seven names config.py imported, and Path in misc.py and logging in
  plugins.py, likewise.
- branch_info in install_plugin was built and never logged; test_config in
  /health was bound and never read (the load_config call is the check).
- An f-string with no placeholders in the asset upload route.
- _installed_plugin_ids wrapped list(manifests.keys()) in try/except;
  _discovered_plugin_manifests always returns a dict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): start.py logs its startup lines and drops unreachable branches

The startup banner went to stdout with print(); it goes through a logger
now, which the app import has already configured, so it reaches the journal
with a level and timestamp like every other line. The "no addresses" branch
is gone: get_local_ips() always returns at least "localhost".

The except around app.run re-raised "only if it's not a client
disconnection error" from inside the branch that had just established it
was one, so that raise could not run. It is one check now, on a named
tuple of the errnos, which the werkzeug log filter uses too. The comment
on threaded=True counts three SSE endpoints, which is how many there are.
Trailing whitespace is stripped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): save_main_config names its General fields once

The General tab's field names were listed twice, once to detect a General
form post and again, with four more, to keep the remaining-keys merge from
storing them as top-level keys. GENERAL_FIELDS and _MAPPED_TOP_LEVEL_FIELDS
hold them now, and the four per-section skip checks are one set.

The comment on that merge said plugin configs are handled "here too", and
"(including plugin keys)". Plugin sections are handled and removed from the
body before it runs; the comment says so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): plugin directories come from the plugin manager only

Six lookups fell back to PROJECT_ROOT/plugins/<id> when there was no plugin
manager: GET /plugins/config's of-the-day data, POST /plugins/action, the
plugin static-file route, the calendar credentials upload and the calendar
OAuth routes. The loader never scans plugins/ (PluginManager.discover_plugins
reads only the configured directory, plugin-repos by default), so what they
found there was a plugin that never runs. _plugin_directory() asks the
manager and answers None without one, which each route already reports as
"not found".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web-backend

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:35:49 -04:00
Chuck 19ee84a760 Merge remote-tracking branch 'origin/main' into claude/hdpi-scroll-performance-antialiasing-4ae609 2026-09-24 17:35:41 -04:00
ChuckandClaude Opus 5.5 9a705c2ca4 revert(vegas): gate only the prefetch thread; gating ESPN fetches measured worse
84043468 also gated the ESPN chunk fetches and the background data
service's workers, for the hourly sports refresh. A burst test on hdpi
(baseball and football refreshing every 5 minutes, 10-minute soaks, G F F G):

  G  prefetch gated only        0.87%, 0.83% late; 6+ late 20, 16; fetches 0.3-1.8s
  F  + fetch threads gated      1.23%, 1.05% late; 6+ late 12, 16; fetches 1.6-4.6s

Every parked fetch thread wakes at each swap and has to take the GIL again
just to park at the end of the window, so twenty of them cost more than
they saved, and the fetches ran two to three times as long. The plugins'
own copies of espn_dates (half the burst) were never gated anyway.

espn_dates and the background data service go back to main's versions and
the module-level active gate goes. Kept from that commit: the render thread
is never gated, a live refresh from another thread can't take its place,
and nested blocks keep the outer boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:35:34 -04:00
ChuckandClaude Opus 5.5 7b90759252 fix: /errors stack traces, Wi-Fi disconnect and save, plugin fonts, API cache TTL (#636)
* fix(errors): record the exception's own stack trace

record_error() called traceback.format_exc(), which only sees an
exception while its except block is running. plugin_executor records
exceptions caught on a worker thread after that block has ended, so
every trace on /errors read "NoneType: None". The trace is now built
from the exception's __traceback__. The executor's log call had the
same problem with exc_info=True and now passes the exception.

record_error() also merged LEDMatrixError context into the caller's
dict in place; it now works on a copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(wifi): point at configure_wifi_permissions.sh instead of a sudoers list

The module docstring told users to grant NOPASSWD sudo on iptables and
ip. configure_wifi_permissions.sh refuses those grants on purpose: a
wildcard rule for either runs an arbitrary program as root. Point at
the script and say why it leaves them out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(wifi): disconnect finds the saved profile by SSID

disconnect_from_network() asked `nmcli -f NAME,802-11-wireless.ssid
connection show` for the profile to take down, but nmcli rejects that
column for `connection show`, so the lookup always failed and only the
device was disconnected. The per-profile lookup _connect_nmcli() already
used is now _find_profile_for_ssid(), and both callers share it. It
also splits terse output on the last colon and unescapes "\:", so a
profile name containing a colon is found.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(wifi): write wifi_config.json atomically and report a failed save

_save_config() opened the file for writing in place and swallowed any
error, so a wifi_config.json left owned by root made the web toggle for
auto-enabling AP mode report success while nothing was saved, and a
crash mid-write could truncate the file. It now uses atomic_write_json,
which also keeps the file's owner and shared group when root saves it,
and returns False on failure. POST /wifi/ap/auto-enable answers 500 in
that case.

The file is now written with indent=4, like the other config files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): resolve plugin:// fonts in the plugin's own directory

FontManager looked for a plugin's bundled fonts under Path("plugins") /
plugin_id: relative to the process cwd, and not the default install
directory (plugin-repos/), so a manifest's plugin:// fonts never loaded.

register_plugin_fonts() takes an optional plugin_dir, and PluginManager
passes the directory it loaded the plugin from. Callers that omit it get
a lookup in the configured plugin_system.plugins_directory, then plugins/,
resolved against the install root.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(api-helper): cache responses for the requested cache_ttl

APIHelper.get(cache_ttl=...) and set_cache(ttl=...) dropped the ttl on
the claim that CacheManager does not support one, but CacheManager.set()
takes a ttl, stores it with the entry, and both cache tiers honour it
over a reader's max_age. Without it every response expired after the
300-second default read age, whatever the plugin asked for. The ttl is
now passed through, and the cache read passes cache_ttl as max_age for
entries written without one. The class docstring describes what the
helper actually does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(style): one scale range for the schema, element_scale and LogoHelper

The generated Scale field allowed 0.1 to 10, element_style's reader
capped at 10 with no floor, and LogoHelper accepted 0.05 to 8 and reset
anything else to 1.0. A logo scale of 9, which the form accepts, drew at
the shipped size.

MIN_ELEMENT_SCALE / MAX_ELEMENT_SCALE (0.1, 10.0) in src.element_style
are now the schema bounds and the clamp every reader applies through
coerce_scale(): a positive number outside the range is clamped, and
anything that is not a finite positive number means the default. That
also stops element_scale() passing NaN through, since min(nan, 10.0)
is nan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(logos): placeholder lands at the requested path; empty logos list

download_missing_logo() wrote its fallback placeholder to
<normalize_abbreviation(abbr)>.png in the logo directory rather than to
the logo_path the caller passed, so it could return True while nothing
existed where the plugin looks (e.g. "TA&M.png" vs "TAANDM.png").
create_placeholder_logo() takes an optional filepath, and
download_missing_logo passes the requested one.

download_missing_logo_for_team() only caught KeyError, so a team whose
"logos" list is empty raised IndexError; it now treats KeyError,
IndexError and TypeError as "no logo URL".

The placeholder is drawn with PLACEHOLDER_SIZE / PLACEHOLDER_BG, the
constants is_placeholder_logo() recognises it by, instead of repeated
literals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): resolve bundled font paths against the install root

TextHelper's default font_dir, the logo placeholder's font and
FontManager's font_overrides.json were all relative to the process cwd,
so a process started anywhere but the install root (the plugin safety
harness, a manual run, a unit without WorkingDirectory) drew with PIL's
default face and read no overrides. They now go through
font_layout.resolve_asset_path; the overrides file sits in the install
root's config/.

The resolver docstrings described an order the code does not follow:
resolve_asset_path never consults the cwd, and sports_shared's
_resolve_font_path tries the cwd first. Both docstrings now say what
the code does, and _resolve_font_path calls resolve_asset_path instead
of probing FontManager for it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sync): the web UI reads the sync status file the display writes

sync_manager writes its status to tempfile.gettempdir(), but
GET /api/v3/sync/status read a hardcoded /tmp/led_matrix_sync_status.json
and defaulted the port to a literal 5765. Wherever TMPDIR is set (or on
any non-/tmp host) the page only ever showed "starting". The endpoint now
uses sync_manager.STATUS_FILE and SYNC_PORT.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(http): the rankings resolver sends the project's User-Agent

DynamicTeamResolver fetched ESPN rankings with a bare requests.get, so
it sent python-requests' default User-Agent, which ESPN rejects; the
AP_TOP_N favourites then resolved to nothing. It now sends
DEFAULT_HTTP_HEADERS. BaseOddsManager carried its own copy of the
User-Agent string and now uses the same shared headers (which also adds
Accept-Language).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(backup): record the core release and read the configured plugin dir

The manifest's ledmatrix_version came from a VERSION file that does not
exist, then from .git/HEAD: a 12-character sha, or "ref: refs/he" when
the branch's ref was packed. It is now src.__version__.

list_installed_plugins() scanned a hardcoded plugin-repos/, so on an
install whose plugin_system.plugins_directory points elsewhere, plugins
missing from plugin_state.json were left out of the backup. It now reads
the configured directory from config/config.json, defaulting to
plugin-repos.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(startup): report a missing display section once

A config without a display section produced three errors for the one
problem ("Missing required configuration key: display", "Display
configuration is missing or empty" and "Display configuration is
missing"), and an empty one produced two. _validate_config now reports
it once, as a missing key or an empty section, and
_validate_display_config leaves it to that.

The module docstring said the validator fails fast; nothing in the
display service calls raise_on_errors(), so it now says the errors are
reported and startup continues.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(wifi): share the copied blocks and name the AP constants

- _parse_nmcli_wifi_list() is the one parser behind _scan_nmcli and
  _scan_nmcli_cached.
- _verify_connected(), _wait_for_device_idle(), _failsafe_ap() and
  _mark_forced() replace blocks that were pasted two or three times in
  the connect and enable-AP paths. The device-idle wait now checks
  before its first one-second sleep instead of after it.
- _check_command() calls _find_command_path() instead of repeating it.
- AP_IP, PORTAL_PORT, AP_PROFILE_NAME and AP_PROFILE_NAMES name values
  that were spelled out 14, 12, 8 and 2 times; the two deletion loops
  now walk the same tuple. The iwconfig status path compares the AP
  address exactly: startswith() also skipped 192.168.4.10-19.
- Dropped a second WIFI.SIGNAL query that repeated the first, a no-op
  "if ssid: continue", the try/except around _connect_wpa_supplicant's
  constant return, and a second save of a scan scan_networks already
  saves.
- _ensure_wifi_radio_enabled's docstring says it returns True when the
  radio state cannot be read at all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(config): drop dead branches and history comments in ConfigManager

- The module docstring pointed plugin authors at update_plugin_config(),
  which does not exist; it now names save_config_atomic() and
  save_raw_file_content().
- load_config's FileNotFoundError handler tested the message for
  "config_secrets.json", but a missing secrets file is handled where it
  is read, so only config.json reaches it; the check is gone.
- save_raw_file_content's `file_type == "main" or "secrets"` guard was
  always true (anything else raised earlier).
- get_raw_file_content('secrets') already returns {} for a missing file,
  so the os.path.exists() in front of two calls to it is gone.
- Comments that narrated earlier behaviour are rewritten as what the
  code does now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(background-data): present-tense comments, drop unused API

- Comments that told the history of each fix (what "used to" happen,
  "the old per-delivery release") now state the invariant the code keeps.
- get_statistics() no longer reports a constant 'queue_size': 0, and the
  uncalled clear_completed_requests() is gone (_cleanup_completed_requests
  does that job on every completion). Neither is referenced in core, the
  web UI or the plugin monorepo.

shutdown_background_service() has no production caller either, but it
is the only way to tear down the get_background_service() singleton,
which the tests rely on, so it stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(odds): drop the unread cache_ttl and merge the odds_data branches

BaseOddsManager loaded base_odds_manager.cache_ttl from config and never
used it: cached odds live for the update interval (get_odds' ttl=interval).
No core or monorepo code reads the attribute, so it is gone along with
its log line. The two consecutive `if odds_data:` blocks are one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(backup): one table for the single-file sections

config, secrets, wifi and ytm_auth were each spelled out in create,
preview, validate and restore. _SINGLE_FILE_SECTIONS lists them once,
with the RestoreOptions flag that restores each, and all four walk it.
Restore error messages keep their wording ("Failed to restore
<file name>").

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(fonts): drop FontManager's write-only state and duplicate logs

- fonts_config, font_metadata and font_dependencies were written and
  never read; the performance_stats keys font_load_times, render_times,
  total_renders and the per-call "resolve" timings
  (_record_performance_metric) likewise. get_performance_stats() reads
  only the counters that remain. Nothing in core or the plugin monorepo
  references any of them.
- A failed BDF load was logged twice, by _load_bdf_font and again by
  get_font; get_font's line is the one kept.
- Removed "NEW:" and commented-out cozette entries, the "Copy font to
  assets/fonts" comment on code that copies nothing, and local imports
  of names the module already imports. The deprecated add_font() now
  resolves assets/fonts against the install root.

The @deprecated methods stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(text-helper): cache loaded fonts; drop the pre-textlength fallback

TextHelper declared _font_cache, cleared it and reported its size, but
never stored anything in it. load_fonts() now keeps each (file, size)
it loads there, so clear_font_cache() and get_font_cache_stats() mean
what they say and repeated load_fonts() calls reuse the fonts.

get_text_width() no longer catches AttributeError for Pillow releases
without ImageDraw.textlength; requirements.txt pins Pillow>=12.2.
The class docstring describes what the helper does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(common): fix wrong docstrings in api_helper, permission_utils, snapshot_policy

- permission_utils called 0o2775 "sticky bit"; the 2 is setgid, which is
  what makes new files take the directory's group.
- snapshot_policy pointed at web_interface/blueprints/api_v3.py, which
  is a package now; the health check is in api_v3/misc.py.
- APIHelper.clear_cache() lost a history note and a fallback to a
  clear() method that neither CacheManager nor the testing
  MockCacheManager has. The session headers are built from
  DEFAULT_HTTP_HEADERS instead of a copy of them, and the module
  docstring says what the module offers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(sports): present-tense comments in the shared scoreboard renderers

- sports_scroll and sports_game_renderer comments that referred to "this
  PR", "the old flat 128px card" or what the renderer "previously" did
  now describe the current behaviour and its reason.
- The block explaining why non-finite settings are rejected sat above
  _score_reserve_width; it describes _center_gap_width and now lives in
  it.
- unshare_element_fonts wrapped its import of font_layout.load_truetype
  in an `except ImportError` that cannot fire inside core; the import
  stays at call time so tests can spy on the pinned loader.
- sports_card docstrings that told the history of a fix say what the
  code does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sports-shared): drop dead code, name the ESPN limit

- _get_weeks_data asked for limit=1000, which fetch_espn_scoreboard
  clamps to ESPN_MAX_LIMIT anyway; it now names that constant. Its
  unused `immediate_events = []` is gone.
- _get_season_schedule_dates() returned ("", "") and has no caller in
  core or the plugin monorepo.
- _should_log keeps its warning_type parameter (part of the inherited
  signature, though nothing in core or the monorepo calls it) and its
  docstring says the cooldown is shared across types.
- An unused ImageFont import is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sync): one follower-mode switch, shared panel defaults

- The class docstring said the leader sends PNG frames. Frames go over
  UDP as raw RGB; PNG is only the Vegas scroll image sent over TCP. It
  now describes both paths.
- _enter_follower_mode() replaces the two copies of "note the leader,
  switch from standalone to follower, log, write status" in the frame
  and scroll-position handlers.
- The rows/cols fallbacks use DEFAULT_ROWS / DEFAULT_COLS from
  src.display_geometry, as chain_length already did.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(style): drop _layout_axis, name the layout group title

- ElementStyleResolver._layout_axis() had no caller in core or the
  plugin monorepo.
- _element_block_from_spec checked spec['size'] was a dict again after
  size_spec already had; it reads size_spec.
- The "Layout Offsets" title written into three generated schema blocks
  is _LAYOUT_TITLE.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(logo-helper): say what the placeholder draws; name the 1.5 box factor

- _create_placeholder_logo's docstring said it draws the team
  abbreviation; it draws an outlined grey box and nothing else. The
  docstring says so, and the "in a real implementation you'd want text"
  comments are gone.
- The 1.5 x panel default logo box, written out six times, is
  DEFAULT_LOGO_BOX_FACTOR.
- ImageDraw is imported with Image at the top of the module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(logos): drop dead code and a duplicate regex in logo_downloader

- _SAFE_LEAGUE_CODE_RE was the same pattern as _SAFE_LEAGUE_RE; both
  checks use the one.
- get_logo_filename_variations reassigned the TA&M case to the list it
  already had; the function returns the two names directly.
- _get_team_name_variations() had no caller in core or the plugin
  monorepo.
- fetch_single_team's docstring was copied from fetch_teams_data; a log
  message read "for{team_id}".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: drop the Pillow<9.1 resample shim and a catch-and-reraise

- adaptive_images fell back to Image.LANCZOS/NEAREST for Pillow < 9.1;
  requirements.txt pins Pillow>=12.2. RESAMPLE_LANCZOS and
  RESAMPLE_NEAREST keep their names (src.common re-exports them).
- CacheManager.save_cache caught CacheError only to re-raise it; the
  disk write is now called directly, with the same result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(api-helper): stop the real CacheManager's cleanup thread

The cache-lifetime tests built a CacheManager and left its cleanup
thread's class-wide claim on the directory in place, which broke
test_cache_cleanup_thread_ownership when it ran later in the session.
The fixture now stops the thread on teardown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): core-common

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:32:29 -04:00
ChuckandClaude Opus 5.5 b11bcfa204 fix(plugins): store and plugin-manager bugs; tidy src/plugin_system (#635)
* fix(store): don't read a ZIP-installed plugin's remote from the LEDMatrix repo

update_plugin looked up remote.origin.url with `git -C <plugin> config
--local` for plugins that are not git checkouts. Under plugin-repos/ git
walks up to the enclosing LEDMatrix repository, so the lookup returned
LEDMatrix's own URL and a plugin missing from the registry was
"reinstalled" from the LEDMatrix repo. Only ask git when the plugin
directory has its own .git, the test _get_local_git_info already uses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(schema): report each missing required field once, by name

validate_config_against_schema ran its own required-fields loop after
Draft7Validator.iter_errors, which already yields one `required` error
per missing field, so every missing top-level field was listed twice.
The validator's copy also printed the schema's whole `required` list
("Missing required property '['api_key', 'city']'") instead of the field.
Drop the loop and take the field name from the error itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): stop mangling repository URLs that contain ".git"

install_from_url and fetch_registry_from_url cleaned URLs with
`rstrip('/').replace('.git', '')`, which removes ".git" anywhere:
https://github.com/user/my.github.io became .../myhub.io, so installing
or browsing that repository asked GitHub for one that does not exist.

Add src/plugin_system/repo_urls.py with one anchored normalize_repo_url(),
same_repo() for comparisons, github_owner_repo() and github_api_headers(),
and use them for the five copies of the owner/repo parsing and GitHub
headers in the store and for saved repositories. GitHub URLs are now
recognised by urlparse().hostname everywhere: _get_latest_commit_info
used a substring test, and _install_from_monorepo_api parsed any host.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): install a repository whose only branch is not main/master

_install_via_git returned None both when every clone failed and when the
last-resort clone of the repository's default branch succeeded.
_install_plugin_impl papered over it with `and not plugin_path.exists()`;
install_from_url did not, so a repository whose only branch is e.g.
`develop` was cloned, then treated as a failure, then "downloaded" from
main/master archives that do not exist.

After a default-branch clone, return the branch the clone checked out
(read from .git/HEAD), so None means failure and nothing else, and give
both callers the same `branch_used is None` fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): judge the memory limit on each call's own growth

monitor_call stores `metrics.memory_mb = max(previous, growth)`, and
_check_limits compared that high-water mark with max_memory_mb. It never
decreases, so once one update() grew the process past the limit every
later call raised ResourceLimitExceeded and the circuit breaker kept
reopening. Pass the call's own RSS growth to _check_limits; keep the
high-water mark for reporting and document what it measures.

Remove ResourceMetrics.update_average_execution_time: nothing called it,
and it overwrote the running total with the average.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): reload_plugin re-reads the manifest from the discovered directory

reload_plugin read `plugins_dir / plugin_id / "manifest.json"`, ignoring
the discovery map and the plugin_dirs rules. For a plugin whose
directory name differs from its manifest id the path did not exist, the
re-read was skipped without a word, and the reload kept the stale
manifest. Resolve the directory with find_plugin_directory, as
load_plugin does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): drop the always-null last_display from plugin state info

PluginStateManager reported `last_display` from `_last_display`, which
nothing ever wrote, so it was null for every plugin. Recording it in
PluginExecutor.execute_display would not help: get_state_info's only
reader is the web process, whose PluginManager never calls display().
Remove the field, its dict and get_last_display() (no caller in core,
the web UI or the plugin monorepo).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(store): share the rollback and requirements helpers, drop dead code

- install_plugin and _reinstall_with_rollback set aside, discard and
  restore the old copy through _set_aside/_discard_backup/_restore_backup
  instead of two copies of the same blocks.
- The loader and the store run the same pre-pip checks through
  contained_plugin_dir() and requirements_to_install() in plugin_loader.
  They still invoke pip differently (sys.executable -m pip vs. the sudo
  wrapper). `except (BrokenPipeError, OSError)` + `isinstance(e, OSError)`
  becomes `except OSError` checking errno.EPIPE.
- load_module never returns None, so load_plugin's check is gone and the
  docstring says what it raises.
- Remove the always-true JSONSCHEMA_AVAILABLE, the inline re-imports of
  re and permission_utils, the fake status_result object nobody reads,
  hasattr(git_error, 'cmd'), a redundant "merge conflict" test and
  `import traceback` (exc_info=True does it).
- Correct comments: install_from_url names the directory for the
  caller's id when given (not always the manifest id), _get_local_git_info
  saves one git subprocess (not four), _enrich calls two helpers,
  search_plugins documents all its arguments, _find_plugin_path states
  its behaviour instead of a TODO, and history narration is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): tidy base_plugin, correct plugin_manager/state comments

- base_plugin: drop the unused `import logging`; get_display_duration
  runs the instance value and the config value through one
  _positive_seconds() helper instead of two copies of the coercion; the
  'static'/'none'/fallback branches of get_vegas_display_mode, which all
  returned FIXED_SEGMENT, are one; fix the mis-indented validate_config
  example; say that get_supported_vegas_modes/get_vegas_segment_width
  are not consulted by core (kept, plugins override them).
- schema_manager: import expand_style_elements normally rather than
  swallowing an ImportError of a core module.
- plugin_manager: the plugins directory is the configured one
  (plugin-repos/ by default), not plugins/; get_config() returns the live
  dict, not a copy, so the interval cache comments say what it saves.
- state_manager: config_version and the file version are not used to
  detect corruption; say what they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): stop writing data/plugin_operations.json

PluginOperationQueue wrote its finished-operation history to
data/plugin_operations.json after every operation, and read it back only
into its own in-memory list, which only get_operation_history() exposes
-- and nothing calls that. The operation-history endpoint reads
OperationHistory (data/operation_history.json). No code in src/,
web_interface/, scripts/ or test/ reads the file.

Drop the history_file/lazy_load parameters and the load/save code; the
bounded in-memory history stays. web_interface/app.py and the
integration test stop passing the removed arguments. An existing
data/plugin_operations.json is left in place (data/* is gitignored).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): plugin-system

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:32:02 -04:00
ChuckandClaude Opus 5.5 3967a6cffc fix(security): re-harden root sudo helpers; installer fixes; ARCHITECTURE and PERMISSIONS docs (#640)
* docs: add ARCHITECTURE and PERMISSIONS guides

ARCHITECTURE.md maps the processes, the state the display and web
services share through the cache, the display loop, the plugin system,
the web UI and the update path, with links into the code and a
where-to-start table.

PERMISSIONS.md lists who owns what after install, both sudoers files
(and why iptables is not granted), the polkit rule, and which
scripts/fix_perms script to run as which user.

Both are linked from the docs index, along with the MQTT bridge README
and src/common/README.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct stale setup, service and troubleshooting claims

- README: quick actions run systemctl on ledmatrix.service (run.py), not
  display_controller.py; use_short_date_format has no effect; the
  installer uses system pip with --break-system-packages, not a venv.
- CONFIG_DEBUGGING: LEDMATRIX_DEBUG must be "true"; logs are in journald.
- GETTING_STARTED, WEB_INTERFACE_GUIDE, TROUBLESHOOTING: enabling a
  plugin, plugin settings, brightness and Vegas settings apply without a
  restart; matrix hardware settings still need one.
- TROUBLESHOOTING: install dependencies with sudo so the root service
  sees them; point permission problems at PERMISSIONS.md instead of a
  project-wide chown.
- ADVANCED_FEATURES: real BackgroundDataService stats keys; Vegas hooks
  return VegasDisplayMode and None falls back to capture; cache files
  are 0660; fix_web_permissions.sh runs as the web user and does not
  touch sudoers.
- STARLARK_APPS_GUIDE: only the linux-arm64 pixlet binary is downloaded.
- HOW_TO_RUN_TESTS: test class examples that exist.
- CLAUDE.md: PluginStoreManager, plugin_dirs.py, monorepo installs via
  the Trees API with ZIP fallback, requirements.txt is optional.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: mark deprecated plugin APIs and state manifest fields once

Methods @deprecated("3.7.0") (the set pinned in test_deprecation.py)
were shown as current API in the quick reference, API reference,
advanced guide, development guide and FONT_MANAGER. Each is now marked
deprecated with its replacement. FONT_MANAGER is rewritten around the
current API; the override editor is gone and override methods are
deprecated.

Required manifest fields were stated three different ways. The API
reference now has one section: the 7 schema-required fields, the 4 the
store refuses without, class_name for the loader, and the 8 to set.
The other guides link to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document every src/common module and every widget

- src/common/README.md covered 7 of 17 modules. It now has a table of
  all of them (purpose, whether plugins import it, release to floor
  on), a short entry each, and logging advice that matches the code.
- SPORTS_UNIFICATION listed two shared modules and called
  sports_helpers the first; it now lists all six.
- The widgets README lists all 28 registered widgets plus the support
  files, and absorbs the parts that only docs/widget-guide.md had
  (x-options.labels, x-advanced, x-display hidden, plugin-file-manager).
  docs/widget-guide.md is now a pointer to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): fix_web_permissions.sh re-hardens the root sudo helpers

The script chowns the whole project to the web user. That included
scripts/fix_perms/safe_plugin_rm.sh and safe_pip_install.sh -- the two
helpers /etc/sudoers.d/ledmatrix_web lets the web user run as root -- so
running it turned both into a root shell for whoever can edit them. It
also re-grouped config_secrets.json away from ledmatrix.

After the chown it now does what first_time_install.sh's Steps 11 and
11.1 do: helpers back to root:root 755, and config_secrets.json back to
the web unit's User=:ledmatrix 640. Each step is non-fatal and prints the
manual command if it fails.

Also fixes what the script and its docs claimed: it never configured
sudoers, its closing hint pointed at ./configure_web_sudo.sh (wrong
path), and the README and ADVANCED_FEATURES.md said to run it with sudo,
which it refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): validate and harden every sudoers drop-in the scripts write

configure_wifi_permissions.sh copied its rules into
/etc/sudoers.d/ledmatrix_wifi without `visudo -c`. A malformed drop-in
makes sudo refuse every command for every user, which on a headless Pi
leaves no way back in. It now checks first and leaves the installed file
alone when the rules do not parse, as the other two writers do. (It
already used mktemp, so that part of the review did not apply.)

It also grants the two literal commands wifi_manager.py runs for
NetworkManager's shared-mode dnsmasq drop-in -- `cp
/tmp/ledmatrix-nm-dnsmasq.conf .../dnsmasq-shared.d/ledmatrix-captive.conf`
and `rm -f` of that file. The directory's mkdir was granted, the file was
not. Both are pinned in test_sudo_allowlist_covers_calls.py.

configure_web_sudo.sh wrote its rules to /tmp/ledmatrix_web_sudoers_$$,
a predictable name in a world-writable directory; it now uses mktemp with
an EXIT trap, as first_time_install.sh does. It sets mode 440 on the
installed file instead of leaving the temp file's mode, and finds visudo
in /usr/sbin when that is not on the user's PATH, which skipped the
check silently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): escape the project path in the DNS-fix and MQTT unit renderers

install_dns_fix.sh and install_mqtt_bridge.sh substituted
__PROJECT_ROOT_DIR__ with the raw path, while the other three renderers
go through sed_escape_replacement from lib_systemd_render.sh. A checkout
under a path containing `&`, `\` or `|` rendered a corrupted unit from
these two only. Both now source the helper and use it, and a test checks
that every placeholder substitution in scripts/install uses an escaped
value.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): stop the installer scripts reporting things that are not true

- first_time_install.sh printed "Password: ledmatrix123" for the setup
  access point. wifi_manager creates it as an open network ("No
  password" on the panel), so it now says so.
- Step 10.1 printed "✓ WiFi management permissions configured" straight
  after its own failure message; install_wifi_monitor.sh printed
  "✓ Package installation completed" after a failed apt install. The
  tick now only follows success.
- Step 7 printed "Web dependencies already installed ... in Step 5" in
  the one branch that runs because Step 5 did not install them, then
  created .web_deps_installed on that basis. It now warns and leaves the
  marker off so the next run retries, as the comment below it intends.
- check_system_compatibility.sh called Debian 12 Bookworm "full
  compatibility confirmed" while first_time_install.sh refuses anything
  but Debian 13. Bookworm, older Debian and non-Debian systems are now
  errors. Its counters used ((X++)), which under `set -e` exits the
  script at the first warning or error (the expression is 0), so the
  check never reached its summary on any system with one.
- configure_web_sudo.sh and configure_wifi_permissions.sh finished by
  testing `sudo -n test -f ...` and `sudo -n nmcli device status`,
  neither of which is granted, so they always reported a failure. They
  now ask `sudo -n -l` about commands the new rules do grant, which
  checks the rule without running anything.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): print the completion summary before rebooting

With -y -- and so for every one-shot `curl | bash` install, which always
passes -y -- first_time_install.sh ran `reboot` about 180 lines before
its "Installation Complete / Web UI Access" summary. reboot returns at
once, so the summary printed while the Pi was going down and the SSH
session usually dropped before the web UI address could be read.

The reboot block moves, unchanged, to the very end of the script. The
interactive prompt now also follows the summary. Because the summary now
runs before the -y reboot, its one command that could fail under
`set -Eeuo pipefail` (the SSID lookup, when nmcli reports a connected
device but no active network line) gets `|| true`; a missing SSID was
already handled as "SSID unknown".

one-shot-install.sh prints its "Next steps" after the installer returns,
by which time the reboot is under way, so it now says so, and README's
Quick Install mentions the automatic reboot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(scripts): correct wrong comments and messages, drop dead code

No behaviour change except the output text noted below.

- 2775 is setgid, not the sticky bit (first_time_install.sh Step 3.1,
  fix_plugin_permissions.sh), and root needs no "PWM hardware access"
  to plugin files.
- The 777 comments in first_time_install.sh Step 3's fallback and
  fix_assets_permissions.sh said root needs it to write. Root ignores
  mode bits; the comments now say what 777 actually opens. The 777
  itself is unchanged.
- apt_remove ends in `|| true`, so Step 12's "Some packages could not be
  removed" branch could never run; it is gone and the helper stays
  non-fatal.
- detect_web_service_user's comment named Step 8 for the web unit
  (install_service.sh installs it in Step 7.5) and now says which
  branch actually runs.
- Step 5 described an "already installed" check that does not exist;
  the ACTUAL_USER comment described the re-exec backwards.
- on_error printed a literal "\n" before "Common fixes:".
- Dead code: one-shot-install.sh's uncalled fix_tmp_permissions,
  LEDMATRIX_ELEVATED=1 (never read) on the sudo re-exec, and
  configure_web_sudo.sh's unused PYTHON_PATH, which also made a missing
  python3 fatal for rules that never mention it.
- start_display.sh / stop_display.sh said "for user: <you>"; the
  service runs as root.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(fix_perms): fix_cache_permissions.sh uses setup_cache.sh's model

There were two models for /var/cache/ledmatrix. setup_cache.sh (the
installer's Step 2) and install_web_service.sh share it through the
ledmatrix group: root:ledmatrix, 2775, files 660, which is also what
DiskCache relies on to give files the directory's group.
fix_cache_permissions.sh instead made it 777 and re-grouped it to the
invoking user's group, undoing that.

It now runs setup_cache.sh for /var/cache/ledmatrix and keeps its own
handling of ~/.ledmatrix_cache. Dropped: /var/cache/ledmatrix/
placeholder_logos (nothing reads it) and the checks against the
`daemon` user (no service runs as daemon).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: pin actions/checkout in the Claude workflows, drop template comments

claude.yml and claude-code-review.yml used actions/checkout@v4 while
test.yml and release-version-check.yml pin the v4.2.2 commit SHA; they
now pin the same SHA. The commented-out starter-template settings
(prompt, claude_args, paths, author filter) are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(scripts): index every script and list removal candidates

New scripts/README.md gives one line per top-level script and scripts
directory, marked keep, dev-only or diagnostic, and lists the eight
scripts nothing in the repo refers to as candidates for removal (kept
for now). The install, utils and dev READMEs now list the files they
were missing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: tighten two checks that mutation testing showed were too loose

- The wifi sudoers check matched `visudo -c -f "$TEMP_SUDOERS"` in the
  error report too, so replacing the check with `if false` still passed.
  It now requires the command as the condition.
- The summary test never had the setup access point up, so reinstating
  the bogus "Password: ledmatrix123" line went unnoticed. A case with
  hostapd active now checks the AP is described as open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(permissions): describe the repaired fix_perms scripts and new WiFi grants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): docs-scripts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:31:41 -04:00
Chuck e39f65fce0 Merge remote-tracking branch 'origin/claude/frame-timing-harness' into claude/offscreen-rendering 2026-09-24 17:02:27 -04:00
ChuckandClaude Opus 5.5 d37a3a712a style(perf): say why render_bench fell back; mark the stats path as a safe fixed name (Codacy)
Codacy (Bandit B110, B108). The bench's silent except now prints why it read
config.json directly. The stats file's fixed name in /dev/shm is safe:
write() goes through mkstemp and os.replace, which replaces a planted
symlink instead of following it; the comment says so and marks it nosec.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:02:15 -04:00
ChuckandClaude Opus 5.5 409b63bc17 fix(vegas): tell code the gate must not park in by module name, not path
The gate never parks a thread inside logging, threading, importlib or the
cache, and it matched those as substrings of each frame's file path. On
GitHub's runners Python lives under /opt/hostedtoolcache, so every stdlib
frame said "cache" and the gate never parked anything -- three tests failed
there and passed here. A virtualenv under ~/.cache would have done the same
on a Pi. Match the frame's module name (f_globals['__name__']) instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 16:55:45 -04:00
ChuckandClaude Opus 5.5 84043468ca feat(vegas): ESPN fetches give way to the render thread too
The hourly sports refresh froze hdpi's Vegas scroll for 1.9s: about twenty
espn-chunk threads fetching and parsing at once, and the render thread
(and the stall watchdog) queued behind all of them for the GIL. The
prefetch gate only covered the prefetch thread.

Vegas now makes its gate the active one (render_gate.set_active), and
render_gate.yielding() gives way through it when there is one and does
nothing otherwise. espn_dates wraps each chunk fetch in it (behind the
same import fallback as json_body, for the copies plugins bundle), and the
background data service wraps each worker. The render thread is never
gated -- the first thread to swap is exempt, and a plugin pushing a live
refresh from its update thread cannot take its place -- and nested blocks
keep the outermost frame as the boundary for the lock checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 16:40:55 -04:00
Chuck a6648ec991 Merge branch 'claude/frame-timing-harness' into claude/offscreen-rendering
# Conflicts:
#	src/display_manager.py
2026-09-24 16:34:25 -04:00
Chuck 1df6def8f7 Merge branch 'claude/hdpi-scroll-performance-antialiasing-4ae609' into claude/offscreen-rendering 2026-09-24 16:34:05 -04:00
Chuck 6279530c16 Merge remote-tracking branch 'origin/main' into claude/hdpi-scroll-performance-antialiasing-4ae609 2026-09-24 16:32:10 -04:00
Chuck 52bc520335 Merge remote-tracking branch 'origin/main' into claude/frame-timing-harness
# Conflicts:
#	CHANGELOG.md
#	docs/SCROLL_PERFORMANCE.md
2026-09-24 16:30:49 -04:00
ChuckandClaude Opus 5.5 1afb2383cd feat(vegas): prefetch_gate on by default, after an A/B/C soak on hdpi
Two runs per arm, about 81,000 frames each, order A B C C B A:

  A  step 1 as is            0.90% late, 20.1 per 10k two+ refreshes late
  B  switch_interval_ms 1    0.78% late, 15.8 per 10k
  C  prefetch_gate           0.60% late,  2.5 per 10k

No freezes in any arm, and the next group was ready at every strip
extension, so parking the prefetch thread (3-6s per 8-minute run) cost
nothing visible. The gate is now on unless vegas_scroll.prefetch_gate is
false; on a stock binding it cannot work and says so at INFO once a run
rather than warning on every install. switch_interval_ms stays off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 16:07:26 -04:00
ChuckandClaude Opus 5.5 4e61d7248a refactor(web): one error-response path for api_v3 (#624)
* refactor(web): answer unhandled api_v3 errors from one blueprint handler

Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.

It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.

Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.

HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.

Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin action errors name the real failure, not UnboundLocalError

execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.

Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop the error category and exception-name code guessing

WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.

from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.

suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.

The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one call for the from_exception error responses

Nine plugin routes built a structured error by hand:

    from src.web_interface.errors import WebInterfaceError
    error = WebInterfaceError.from_exception(e, ErrorCode.X)
    return error_response(error.error_code, error.message,
                          details=error.details, context=error.context,
                          status_code=500)

That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): one api_v3 error-response path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:53:19 -04:00
ChuckandClaude Opus 5.5 ece416c4e5 refactor(plugins): one plugin-directory resolver (#623)
* refactor(plugins): one resolver for plugin id -> directory

Five places mapped a plugin id to its directory, each with its own rules
and each re-reading manifests per lookup: PluginManager discovery and
get_plugin_directory, PluginLoader.find_plugin_directory,
PluginStoreManager._find_plugin_path / list_installed_plugins, and
state_reconciliation.disk_plugin_ids. They disagreed on backup dirs,
on whether the manifest id or the directory name is the id, on duplicate
ids and on path safety.

src/plugin_system/plugin_dirs.py now holds the rules once:
PluginDirectoryIndex scans one directory and reads each manifest once;
resolve_plugin_dir() searches directories in order. What legitimately
differs per caller is an explicit argument: search dirs (discovery and
the loader: configured dir only; the store: configured then sibling
plugins/), ledmatrix- prefix (not for the store), case folding (loader
only), manifest pass (not for get_plugin_directory, whose discovery map
already holds it).

Behaviour changes, all for layouts installs do not produce:
- a directory whose manifest declares the id beats one merely named for
  it (discovery already worked this way; the loader and store now agree)
- the store searches the configured dir completely before plugins/
- backup and hidden dirs are skipped everywhere (the loader's case and
  manifest scans and list_installed_plugins used to return them)
- duplicate ids resolve deterministically (exact name, then
  ledmatrix-<id>, then by name) with a one-time warning; discovery no
  longer lists the id twice
- disk_plugin_ids / list_installed_plugins report manifest ids, falling
  back to the directory name; auto-update looks the directory up
- ids that are not one plain path segment resolve to nothing in every
  caller (the loader used to truncate them, the store to join them)

The .standalone-backup- marker is one constant, BACKUP_MARKER, used by
store_manager's rename-aside names and every lookup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): one plugin-directory resolver

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:52:52 -04:00
ChuckandClaude Opus 5.5 13bbb537f3 refactor(web): one logging setup and one TTL cache for the web process (#621)
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG

The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.

app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.

Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.

The duplicate module is deleted; nothing else imported it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one thread-safe TTL cache for the web process

web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.

Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
  counted and get_cached defaulted to 60s. An entry now expires after the TTL
  it was stored with; a reader's ttl_seconds can only shorten that. Both
  current callers pass the same value on both sides (fonts_catalog 300s,
  system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
  same expired key could raise KeyError (reproduced), which the endpoints
  turned into a 500.

app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.

Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web logging and TTL cache

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): only ask systemctl about known units

Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): response_time_ms reads the same clock request_logging stamps

request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:52:17 -04:00
ChuckandClaude Opus 5.5 afe9001aed refactor(fonts): one BDF loader and one BDF rasterizer (#627)
* refactor(fonts): one BDF loader and one BDF rasterizer

BDF faces were loaded three ways (FontManager._load_bdf_font,
element_style._load_bdf, DisplayManager._load_fonts) and drawn by two
copies of the same per-pixel loop (DisplayManager._draw_bdf_text and the
plugin test harness's "replicated" copy), which golden images and
check_plugin/dev_server previews rely on matching the panel.

src/common/bdf_font.py now owns both:
- load_bdf_face(path, size) -> (face, realised_px): native-strike fallback
  for sizes the file lacks, one bounded LRU cache keyed on path, size and
  mtime. FontManager, element_style and DisplayManager delegate to it;
  read_bdf_native_size moves here (the old names delegate).
- draw_bdf_text(draw, text, x, y, face, color, clip): builds each glyph as
  a 1-bit mask and fills it with ImageDraw.bitmap instead of a draw.point
  per pixel. A blending Draw (RGB image, "RGBA" mode) keeps the point path
  so translucent colours still blend.

Pixel-identical: 220,032 renders (every bundled BDF at native and
off-strike sizes, 14 strings, 4 colours, clipped on every edge, through
each old loader x rasterizer) match origin/main byte for byte.
test/test_bdf_font.py keeps a lightweight version against a frozen copy of
the old loop. DisplayManager._draw_bdf_text goes from 1.4-23 ms to about
0.1 ms per string (the old loop re-read FreeType's buffer as a Python list
for every pixel).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(testing): harness calendar_font is sized like the panel's

VisualTestDisplayManager built its 5x7 calendar_font / bdf_5x7_font as a
bare freetype.Face. With no size set its ascender reads 0, so BDF text
drawn with it landed 6px above where DisplayManager draws it -- entirely
off the canvas at y=0 -- and get_font_height() returned 0. Golden images
and check_plugin / dev_server previews showed text the panel does not.

Load it through load_bdf_face at the panel's 7px, so it is the very face
DisplayManager uses. Across the differential run this changes only the
cases drawn with the harness's own calendar_font (968 of 220,032), which
now match the panel's output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): one BDF face per thread

The shared face cache now hands every loader (FontManager, element_style,
DisplayManager, the harness) the same freetype.Face. FreeType does not allow
two threads to use one face at once, since load_char rewrites its glyph
slot, so key the cache by thread as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:53 -04:00
ChuckandClaude Opus 5.5 abedc46104 refactor(sports): merge the sports_shared/sports_card twins that behave identically (#626)
* refactor(sports): wrap the sports_card twins that behave identically

SportsCoreSharedMixin (switch mode, via each scoreboard's sports.py) and
sports_card (scroll/Vegas mode, via game_renderer.py) carried the same
helpers twice. test/test_sports_twins.py now calls every pair with the
same inputs -- the eight scoreboards' harness fixture games in flat,
flat+nested and nested-only shapes, plus edge cases (favourites by id and
abbreviation, NRL's colliding abbreviations, missing and non-numeric
scores, bad zones, out-of-range dates, shared font faces).

Identical pairs become thin wrappers over the sports_card function:
_card_option, _vs_text, _format_game_time, _coerce_rgb, _crisp_size (with
the class's own tables), _unshare_element_fonts (with the class's own
element map, via a new optional argument), and the colour/month/weekday/
font-grid tables (dicts copied, not aliased). _format_game_date shares the
card's formatting body but keeps its own setting, weekday zone and month
table; _schema_font_size shares the parser but keeps its per-class cache,
because a reloaded plugin gets new classes and a shared path cache would
stop it seeing an edited schema. _resolve_font_size agrees but keeps its
body so it still dispatches through the overridable hooks.

No behaviour change: old and new mixin/card agree on all 22,994
comparisons over the test corpus, and the pairs that do differ
(favourite-result colours on nested payloads and by favourites source,
the weekday's timezone, the element-name map, per-mode colours) are left
alone and pinned in TestPinnedDivergence for an owner decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(sports): pin that an ambiguous NRL abbreviation tints in both modes

NRL's resolver passes a shared abbreviation ("NEW") through with an error
and its _is_favorite_game matches ids only, but both favourite-colour
helpers match on abbreviation as well, so both display modes tint a
Knights or Warriors result for a user who typed "NEW". The twins agree;
neither consults the _favorite_key seam. Pinned so a fix is deliberate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:35 -04:00
ChuckandClaude Opus 5.5 1fe7237799 refactor(install): generate the web sudoers rules in one place (#622)
* refactor(install): generate the web sudoers rules in one place

/etc/sudoers.d/ledmatrix_web was written by two copies of the same
allow-list: a heredoc in first_time_install.sh Step 10 and a block of
echo lines in scripts/install/configure_web_sudo.sh. They drifted before
(safe_pip_install.sh was granted by one only), and a test existed just
to catch that.

Both now call web_sudoers_rules() from the new
scripts/install/lib_sudoers.sh and keep their own validate (visudo -c),
install and confirm flows.

- first_time_install.sh output is byte-for-byte unchanged, so a device
  re-running the installer gets "already up to date". If the library is
  missing, Step 10 keeps the installed file and carries on, the same way
  it handles rules that fail visudo (an empty file would pass visudo).
- configure_web_sudo.sh now writes the installer's layout: same 18 rules,
  different comments and order. It still leaves out reboot, poweroff and
  journalctl when they are missing; the library does that for both.

The drift test now pins the generator's grants, checks that neither
installer writes rules of its own, and runs each installer's call line
to check the argument order. Tests that read the rule text now read the
library.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(install): detect the web service user in one function

first_time_install.sh pasted the same WEB_SERVICE_USER detection block
three times (Step 3.1's fallback, the plugin-repos setup and Step 11).
The copies were identical apart from comments; they now call
detect_web_service_user(), whose body is that block unchanged.

Behaviour is the same: the function sets the same global and always
returns 0, as the inline if-chain did. Checked on Linux against all
three original copies across 13 layouts (installed unit with and without
User=, the repo as shipped, each grep branch, template placeholders).

The comment notes that the install_web_service.sh / install_service.sh
greps no longer match anything, so until Step 8 installs the unit the
result is "root". That behaviour is left as it was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:14 -04:00
ChuckandClaude Opus 5.5 ddf5f085a5 perf(cache): tell a stale record from its header instead of parsing it (#633)
The sports plugins cache whole season schedules: 53MB for MLB, 18MB for
NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file
takes ~1.8s with the GIL held, and every thread in the display service
waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with
the interpreter itself blocked, right on these reads. When a season record
expired, DiskCache.get paid that whole parse only to find the timestamp
too old and throw the result away.

CacheManager.set now writes timestamp and ttl ahead of the data, and
DiskCache.get reads them from the first 256 bytes of the file, applying
the same rule as before (a per-entry ttl wins over max_age; no limit
means never stale). A record that is stale is refused without being
parsed. Files in the old layout, and records from other writers, don't
match the header and are parsed in full as before.

Also: ESPN responses in the background data service and espn_dates are
parsed with orjson when it is installed (src/common/json_body.py). The
stdlib parser behind response.json() takes 3.1s on the MLB season
against orjson's 1.8s, both with the GIL held. espn_dates imports it with
a fallback, since plugins bundle copies of that module for older cores.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:50:50 -04:00
ChuckandClaude Opus 5.5 5baf983fe0 docs(scroll): explain the tear across the middle on fast scrolls (#620)
* docs(scroll): explain the tear across the middle on fast scrolls

A 1:32-multiplexed 64-row panel lights row 31 almost a whole refresh after
row 32, so fast scrolls show a sideways offset at mid-height of about
speed x refresh period. Documents the cause, how to read the real refresh
rate (show_refresh_rate prints with a carriage return), what was measured on
a single-chain 2x128x64 Pi 4 (pwm_bits, gpio_slowdown and an uncapped
refresh barely help; gpio_slowdown 2 glitches), and the fix that does help:
fewer pixels per output via parallel chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(scroll): limit the 1:32 row-pair explanation to panels that scan that way

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:50:33 -04:00
ChuckandClaude Opus 5.5 82f3a3a3e4 fix(redaction): make credential redaction linear, not quadratic (#631)
* fix(redaction): make URL-userinfo redaction linear, not quadratic

_REDACT_URL_USERINFO could start a match at every letter of a run of
scheme characters, and each attempt read to the end of the run looking
for `://`. On a long unbroken run of letters or digits (a hex digest, an
ID, part of a response body) that is quadratic: 1.6s for 20k characters.

The display service redacts every message, stack trace and context value
it publishes in the error snapshot, holding the aggregator lock, and
re.sub holds the GIL for the whole call, so one such exception stalled
every thread, render loop included (~0.5s measured for 20k chars of hex).
It also made test_snapshot_stays_small the slowest test in the suite by
far: 142s of a 383s run, 139s of it in this one regex.

A match may now only start where a run of scheme characters starts
(negative lookbehind). Leading digits and `+.-` are captured in group 1
so the substitution restores them, and the scheme still has to start
with a letter, so what gets redacted is unchanged: old and new output
were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k
~6ms, and test_snapshot_stays_small takes 0.8s.

test/test_redaction.py pins the exact output for schemes that begin after
digits or `+.-`, and bounds 50k-character runs at 1s; against the old
pattern those timing tests fail at 3-11s each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK

* fix(redaction): make Authorization-header redaction linear too

_REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two
`\s*` separated only by an optional quote. With no quote, a whitespace
run could be split between them in every possible way, and when no
credential followed (end of text, or `,` `"` `<` ...) the engine tried
them all before giving up: quadratic, 8s for `authorization:` and 20k
spaces, 17s with `Proxy-Authorization:` (tried again at the inner
`authorization`). Same stall as the URL pattern: re.sub holds the GIL,
and the display service redacts everything it publishes.

The quote and the whitespace after it are now one optional unit,
`\s*(?:["\']\s*)?`, which matches the same strings with only one way to
split them. Output is identical to the old pattern on 300k fuzzed
inputs; 20k spaces now take ~1.6ms. A scan of all three redaction
patterns over prefix/run/suffix shapes finds none left that scales
superlinearly.

test/test_redaction.py pins exact output for quoted, tabbed, multi-line
and credential-less headers, and bounds header + 20k whitespace at 1s;
against the previous pattern those fail at 8-17s each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:49:51 -04:00
ChuckandClaude Opus 5.5 c1ce0b7b04 fix(web): two api_v3 paths called names that no longer exist (#625)
The Pixlet editor stop route restarts the display after a SIGKILL with
_run_systemctl_command, which starlark.py never imported (since #554). The
Starlark device-location resolver fell back to _ensure_cache_manager, which
#609 deleted; the resolver already accepts no cache manager. Both raised
NameError on the rare path that reaches them. pyflakes finds no other
undefined names in src/ or web_interface/.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:49:29 -04:00
ChuckandClaude Opus 5.5 14c38a3189 fix(perf): count a stall even when the scroll state went missing across it
On hdpi the stall watchdog logged a 1.9s stall during the hourly sports
refresh that the soak report never had: its worst gap was 655ms. The frame
that ended the stall was recorded as static, so its interval was dropped.

"Scrolling" is DisplayManager's scroll state at the moment a frame is
presented, and it goes missing mid-scroll: it expires after 2s without
activity, and any thread can clear it. Plugins call
set_scrolling_state(False) from their own display() (news, stocks, the odds
ticker's fallback), and Vegas captures some of those on the render thread
between two of its own frames. Vegas sets the state again only after its
next frame, so that frame is recorded as static -- along with the capture
or stall it followed.

One static frame between two scrolling frames, with the scroll resuming
within RESUME_SECONDS (1s), is now a frame of the scroll and both of its
intervals count, the first at the scroll's own hold (clearing the state
drops the hold to 1 too). Two static frames in a row still end the scroll.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:23:18 -04:00
ChuckandClaude Opus 5.5 de54fc879a docs(offscreen): the two GIL experiments, and what each can and cannot cover
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 14:27:37 -04:00
ChuckandClaude Opus 5.5 b67818d5c5 feat(perf): LEDMATRIX_STALL_WATCHDOG_MS lowers the stall watchdog's threshold
250ms catches freezes; the hitches left on hdpi are frames 2-5 refreshes
late, which look like the render thread waiting for the GIL. At 30ms the
watchdog dumps those too, naming what the other threads were running when
the frame missed. It polls at a third of the threshold so a stall one poll
long is still seen, which costs some GIL time of its own: a diagnostic
setting, not one to soak with.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 14:26:04 -04:00
ChuckandClaude Opus 5.5 4f30bc62ca experiment(vegas): vegas_scroll.prefetch_gate runs the prefetch only while the render thread waits on vsync
The render thread spends most of each refresh in SwapOnVSync with the GIL
released, then needs it back the moment the swap returns. With plugin
rendering on the prefetch thread, it often has to wait for it -- behind
bytecode for up to the switch interval, behind a GIL-holding C call for as
long as that takes -- and hdpi's late frames of 2-5 refreshes went up.

src/common/render_gate.py opens a window around each swap, up to just
before the refresh the swap will return on, and a profile hook on the
prefetch thread parks it outside that window. It is never parked holding a
lock the render thread also takes (the Vegas buffer, cache and state
locks, logging, threading, importlib, the cache), never when no frame has
been swapped for 50ms, and never for more than 50ms at a time.

Off by default and ignored on a binding that keeps the GIL in
SwapOnVSync, where the window would never let the prefetch run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 14:23:09 -04:00
ChuckandClaude Opus 5.5 55a2760892 experiment(vegas): vegas_scroll.switch_interval_ms shortens the GIL switch interval during a run
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 14:13:29 -04:00
ChuckandClaude Opus 5.5 f79618d4f7 refactor(bench): grade render_bench with the shared frame-timing recorder
render_bench.py (from the parallel perf/render-bench work) had its own
grading module, frame_pacing, with its own definition of a missed frame
and its own refresh estimate. The soak already had both in frame_timing,
so the two could have drifted apart on what "late" means.

The bench now gives the display manager a fresh FrameTimingRecorder,
drains it synchronously at the start and end of the graded run, and prints
frame_soak's report with frame_soak's verdict. Its workload is unchanged:
the synthetic strip, --busy load, the shared speed resolver, the
per-frame scrolling announcement. frame_pacing, its tests and its
src.common exports are removed; measure_refresh_hz moves to frame_timing,
where scroll_speeds.py now finds it.

Two ideas from frame_pacing carry over. The bench seeds the recorder with
the idle refresh it measures, so a loop that free-runs (the 827fps bug
the first bench caught) shows as early frames and one stuck at half rate
as late frames, where an estimate taken from their own intervals finds
both self-consistent. And the soak, which has no idle measurement, now
calls a run NOT LOCKED when its refresh estimate beats the configured cap.
The report also gives the rate held while rendering.

Docs: the bench becomes "Without the service" under "Soaking a rig",
keeping its hdpi numbers and the idle-vs-rendering refresh finding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 11:56:44 -04:00
ChuckBuildsandClaude Opus 5 d56ec2ab3a feat(bench): measure a rig against the refresh it actually holds
There was no way to answer "does this hardware present every frame on
time?" other than watching the panel. `scripts/render_bench.py` drives the
production path -- a real DisplayManager and ScrollHelper, configured
through the same `scroll_config` resolver every ticker uses -- and grades
the run with a new `src.common.frame_pacing`, exiting non-zero when more
than 0.1% of frames slipped a refresh. Exit 2 when the run could not be set
up at all, so a rig that was never measured cannot pass by accident.

A missed frame is defined exactly: an interval that rounds up to at least
one more refresh than its frame hold asked for. The half-refresh rounding
boundary keeps a frame that ran 1ms long on a 10ms refresh out of the
count, because it still presented on the refresh it was meant to.

The verdict that matters more is NOT LOCKED. A loop that never blocked on
vsync reports a perfect zero misses while presenting nothing -- 8ms frames
on a 100Hz panel all land in the one-refresh bucket while running 25% too
fast -- so the report also checks the typical frame is not shorter than the
panel could physically present. That is what caught the first version of
this benchmark announcing its scrolling state once instead of per frame:
the state expires on an inactivity threshold, the dirty-tracking skip then
fires mid-scroll, and the loop free-ran at 827fps.

And the refresh is read back out of the frames rather than taken from an
idle measurement. Driving the matrix is bit-banging on the same machine, so
pushing frames slows the refresh: a Pi 4 on 512x64 measures 100.4Hz idle
and holds 96.3Hz while scrolling. Both are real, and grading against the
idle figure reports a locked loop as 4% slow -- or, once the gap passes
half a refresh, as missing every frame. The gap between the two is itself
worth watching: a rise in it is a render-cost regression even when nothing
is missed.

Measured on hdpi (Pi 4, 512x64, pwm_bits 8), two minutes each:

  plain      95.44 fps, 8 missed of 11,449 (0.070%)  PASS
  --busy 2   95.41 fps, 3 missed of 11,445 (0.026%)  PASS

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
2026-09-24 11:08:43 -04:00
ChuckandClaude Opus 5.5 8163104581 style(offscreen): lint fixes for the new code (Codacy)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 11:02:22 -04:00
ChuckandClaude Opus 5.5 c883a2fd1e feat(perf): a stall watchdog that logs what the render thread is waiting on
The recorder counts freezes; it cannot say why. hdpi showed 1-2s freezes
in both the #628 and offscreen builds, one lining up with hockey's 2s
NHL fetch on the update thread, and nothing in the logs explained it.

StallWatchdog polls every 50ms from its own thread. When a scroll's last
frame is more than 250ms old (and a scroll is still running, so the end
of a scroll is not a stall), it logs the stack of the thread that
presented that frame and the top of every other thread's, then the
stall's length when frames resume. It also measures how late its own
wake-up was: if it was held up as long as the render thread, the whole
interpreter was blocked (C code holding the GIL), not one thread on a
lock. One dump per 30s at most; LEDMATRIX_STALL_WATCHDOG=0 disables it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 10:58:27 -04:00
ChuckandClaude Opus 5.5 0f68fbcfdc docs(offscreen): step 1 status and first hdpi soak
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 10:03:20 -04:00
ChuckandClaude Opus 5.5 13b5264d11 feat(vegas): render every plugin's ticker content off the render thread
The plugin-facing canvas (DisplayManager.image, draw, matrix) was one
shared object, so any plugin whose Vegas content needed it -- display
capture, scroll-content generation, narrowed rendering -- was deferred to
the render thread and fetched there one at a time. On hdpi that is most
plugins, and each fetch stalled the scroll: news ~320ms, hockey ~660ms,
in bursts whenever the strip extended.

DisplayManager.offscreen() gives the calling thread a canvas of its own.
image, draw and matrix are now properties that resolve to the thread's
surface while it is inside the block and to the shared canvas otherwise,
so the ~100 existing uses become thread-correct unchanged. Inside,
update_display(), the hardware half of clear(), and set_scrolling_state()/
set_frame_hold() are inert, so a plugin drawn for Vegas can neither reach
the panel nor re-pace the live scroll. render_size() is rebuilt on it.
capture_mode() now restores the previous state instead of clearing it,
so it cannot end suppression inside an offscreen block.

The adapter draws every path on its own canvas (_isolated_canvas) and
drops the copy-and-restore of the shared image, which from a background
thread would have written a stale frame back over the render loop's.
Background fetches take the plugin's update/display lock, waiting up to
2s for a running update() and skipping the plugin that round otherwise;
Vegas never took that lock, so render-thread captures already raced
update(). A background fetch that comes back empty is no longer queued
for the render thread.

vegas_scroll.offscreen_prefetch (default true) restores the old deferred
path when false. See docs/OFFSCREEN_RENDERING.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 09:31:27 -04:00
ChuckandClaude Opus 5.5 430e2312f8 docs(offscreen): redraw on real updates with a 10s floor; sync by operation replay
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 09:01:06 -04:00
ChuckandClaude Opus 5.5 1af5d5fe45 docs(offscreen): keep live content fresh: refresh at the gate, replace ahead, patch on screen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 09:01:06 -04:00
ChuckandClaude Opus 5.5 5edc9195ca docs: propose per-thread offscreen rendering for Vegas content
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 09:01:06 -04:00