Commit Graph
8 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5.5 9a1f94f793 fix(starlark): blank app locations use the device location, not San Francisco (#617)
* fix(starlark): blank app locations use the device location, not San Francisco

A Starlark (Tidbyt) app whose Location field is blank rendered at its
author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with
the device city set under General settings. A user in Charlotte, NC got San
Francisco weather and radar with nothing in config.json to explain it.

src/device_location.py fills unset location fields at render time (display
plugin and the web standalone render): the device city is geocoded once via
Open-Meteo, preferring a match in the configured state/country, and cached
permanently. A saved location always wins; if the lookup fails the field is
dropped so the app uses its own default, and the failure is not retried for
30 minutes.

Also fixes the config form: clearing a location omitted the key, and the
save merges, so the old value could never be removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(starlark): say what happens when the device location can't be used

A blank app Location only renders at the device's city when one is set and
the Open-Meteo lookup finds it. With no city, no match, or the geocoder
unreachable (retried after 30 minutes), the app gets no location and keeps
its author's default. The guide, the config page hint, CONFIG_REFERENCE and
the CHANGELOG entry now say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:48:02 -04:00
ChuckandClaude Opus 5 4fe3cdd906 fix(starlark): stop the root display service locking the web UI out of starlark-apps (#604)
* fix(starlark): stop the root display service locking the web UI out

Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with

    sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps

starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.

The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.

The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.

Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.

Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): address the review on the ownership repair

Findings from the automated review of #604.

Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.

install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.

The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.

Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.

NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.

Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.

Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 16:33:12 -04:00
ChuckandClaude Opus 5 116abb0daa fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service

BackgroundDataService always sent a season range first and, on a 400,
fell back to chunks without recording the rejection, so every background
season fetch spent a doomed request and live scoreboards learned nothing
from it (or it from them). The worker now consults and sets the same
6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight
to month/day chunks, and if every chunk fails the range is asked once for
a real error without re-spending the chunks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep plugin asset and action routes inside their directories

POST /plugins/assets/upload, GET /plugins/assets/list and POST
/plugins/assets/delete joined the request's plugin_id onto assets/plugins
unchecked, so '../../config' created, wrote, listed and deleted outside
it. #561 guarded only the route that serves the files. All three now go
through path_safety.resolve_under and answer 400 for anything but a
plain name, and delete only unlinks a metadata path that resolves into
that plugin's uploads directory.

PluginManager.get_plugin_directory refuses ids that are not one plain
path segment, so /plugins/action (which runs a manifest script from the
returned directory) and every other caller get the guard; the action
route also rejects such ids up front, covering its no-manager fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): report a no-op plugin update as already up to date

update_plugin() returns True both for a real update and for "nothing to
do" (a ZIP-installed monorepo plugin already at the registry version, a
bundled plugin). With no git commit to compare, POST /plugins/update
called every such success "updated successfully", so Check & Update All
counted most official plugins as updated on every run.

The route now reads what changed off the plugin itself (commit, else
manifest version, else last_updated) and returns data.update_status
(updated / up_to_date / local_only). The update-all toast is summarised
by PluginInstallManager.summarizeUpdateResults from that status, falling
back to the message for older servers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): scoreboard scroll speed no longer follows target_fps

sports_scroll computed the crisp speed ladder against the global
target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since
frame-locked presentation (#545) the helper steps a fixed number of whole
pixels per presented frame and the panel presents at its real refresh, so
the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a
50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s.

The ladder now uses the display manager's refresh_hz, then
display.hardware.limit_refresh_rate_hz, then the default. target_fps is
not consulted. Docstrings now say scroll_delay is ignored for pacing (no
behaviour change there) and describe the fixed-step model.

Tests: replace the tests that pinned target_fps as the ladder refresh and
described time-based stepping; assert speed independence from target_fps
(unit and end-to-end presented px/s against the real helper), that the
fixed per-frame step is applied, and that scroll_delay does not change
speed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): escape registry and upload values in plugin manager inline handlers

The store, saved-repository and custom-registry buttons built
onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone,
so a custom registry entry whose id contained ' closed the attribute and
added its own handler. One helper, jsStringAttr(), now HTML-escapes the
JSON literal for every one of those handlers, and the store View button
opens only http(s) repo links.

The live window.updateImageList (plugins_manager.js loads last, so its
copy wins over the file-upload widget's) wrote the uploaded file's
original name, path and ids into markup raw; they are escaped now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note plugin asset, action and inline handler guards

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): let the root pip wrapper install web_interface/requirements.txt

Update Code, the automatic update's health check and Install Base
Requirements install web_interface/requirements.txt through
safe_pip_install.sh, which only allowed the root requirements.txt. The
first commit changing that file would fail its dependency install, and
the automatic updater rolls back any update whose dependencies did not
install -- on every device, for every newer commit.

The wrapper now lists both core requirement files. Only their folders
are resolved, so a requirements.txt symlinked out of the project is
compared by its target and refused (previously the root file's own
symlink target was what got allowed). The updater's file list is a
named constant, and a test runs the real wrapper (pip stubbed) on
every file Update Code and the rollback install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): do not retry plugin requests that got an HTTP answer

PluginAPI.request wrapped everything that was not a structured error as
NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a
JSON error without error_code included. Check & Update All retries
NETWORK_ERROR, so those updates were re-sent five more times with
backoff, contrary to the #587 contract that an HTTP error response is
the server's answer.

NETWORK_ERROR now means only that fetch() rejected. Any HTTP response
without an error_code, or with a body that is not JSON, is API_ERROR
with the HTTP status attached. Tested against the shipped api_client.js.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): restart the stats window when an idle gap is dropped by size

#582 dropped an idle gap from the frame stats two ways: the reset_scroll()
sentinel, which also restarts the 5s window timer, and a size guard for
scrollers that never call reset_scroll(), which did not. On that path the
first real frame after the gap found the boundary overdue and logged a
stats line for a one-frame window. Both paths now share one seeding helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): leave plugins alone when update_core's own rollback fails

update_core returns rollback_failed directly when a partial pull or an
update whose health check never started cannot be rolled back. run()
only held plugins back for 'verifying', so those devices still got new
plugin versions and a display restart on top of a core in an unknown
state -- the opposite of what the health-check path does, and of the
3.4.0 changelog (plugins are left alone if the rollback fails).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(api): make the REST reference match the api_v3 package

Every documented request body, query parameter and response shape was
re-checked against the handlers in web_interface/blueprints/api_v3/.
Fixes calls that failed as documented (repo_url, action_id/params,
files/image_id, font_file+font_family, ?font=, cache key,
auto_enable_ap_mode, plugin limit keys), removes the font-override
endpoints dropped in #566, corrects response shapes (plugins/config,
plugins/schema, health, metrics, operation history, github-status,
fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds
the 26 routes it omitted (backup, system auto-update/git, wifi radio,
starlark editor, MQTT bridge, status endpoints, skins).

Documents the merge semantics of partial JSON saves to /config/main and
/plugins/config and the dim-schedule POST accepting GET's days shape,
which land in the same change set. Replaces app.py line numbers and the
removed api_v3.py path with file and function names.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): remove the General-tab plugin system toggles that did nothing

plugin_system.auto_discover, auto_load_enabled and development_mode had
General-tab toggles whose help tips promised dormant plugins and verbose
logging, but nothing reads them: every enabled plugin is discovered and
loaded regardless. Remove the three toggles.

The keys stay tolerated in stored configs. The save handler now stores
a flag only when a client sends it; treating a missing key as an
unchecked box would otherwise rewrite all three to false on every
General-tab save, which still posts plugins_directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(scroll): remove dead code left by #523/#570

- Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read
  them since the numpy blend replaced the scipy path.
- Drop ScrollHelper._last_integer_position and frame_time_target, which
  were written but never read.
- Keep target_fps and set_target_fps() but document them as
  informational: nothing paces off them, yet ledmatrix-elections'
  test_scroll_pacing.py reads helper.target_fps back and third-party
  plugins may call the setter.
- Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to
  throttle", set_sub_pixel_scrolling's "default: True", and
  set_frame_based_scrolling's claim that it steps.

The plugins monorepo was grepped for every removed name; none is used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI

FONT_MANAGER.md told plugins to read display_manager.font_manager, which
does not exist, so a plugin following it failed to load with
AttributeError. The shared FontManager lives on the PluginManager and
BasePlugin._get_font_manager() returns it (with a fallback for harnesses).

Also removes the Fonts-tab override workflow and element-override panels
that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls

/plugins/store/search does not exist (404) and the list endpoint reads
query, not q. The registry guide's curl examples omitted the JSON
Content-Type, so the handlers saw an empty body and answered 400.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): use the shared core-key list in the last three private copies

StartupValidator warned "Plugin 'auto_update' is enabled but not found" on
every display start with auto-update or a dim schedule on; the reserved
plugin-id check missed auto_update, sync, location and the rest; and
ConfigManager's (uncalled) orphan cleanup would have deleted display,
schedule and auto_update. All three now read src/core_config_keys.py, which
also gains CORE_SECRETS_KEYS for the github/youtube secrets sections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): partial JSON saves to /config/main change only what they send

A JSON body with one field reset every checkbox in the sections it touched:
the MQTT bridge's brightness slider turned off disable_hardware_pulsing,
inverse_colors, show_refresh_rate and use_short_date_format, and a
timezone-only save turned off web-UI autostart and weekly auto-updates.
Missing-means-unchecked now applies only to form posts: form-encoded bodies
and the v3 forms, which mark themselves with a hidden __form_section input.

Also on the config routes:
- vegas_min/max_cycle_duration no longer match the generic *_duration rule,
  so they stop landing in display_durations and a blank one no longer
  rejects the whole Display save;
- saving from the Raw JSON editor calls start_setup_if_needed like the
  General form, so enabling auto-update there finishes its setup;
- the schedule and dim-schedule POSTs accept the per-day days.<day> shape
  their GETs return, as well as the flat form keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): install plugin dependencies from the configured plugins directory

install_plugin_dependencies.sh scanned only plugins/, but the Plugin
Store installs into plugin_system.plugins_directory (default
plugin-repos), so the documented "Recommended" fix found 0 plugins on
every store install. It now reads plugins_directory from
config/config.json (relative to the project root or absolute, default
plugin-repos) and also scans plugins/ for dev symlinks, installing a
plugin reached through both only once.

With set -e alone, `pip ... | tee` took tee's exit status, so a failed
pip install was reported as success; set -o pipefail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: replace stale API names, line numbers and the api_v3.py path

- ADVANCED_FEATURES: StreamManager methods that exist
  (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the
  real on-demand status envelope ({status, data: {state, service}})
- app.py:199 / :144 / :607-619 line citations and
  web_interface/blueprints/api_v3.py (now a package) replaced with file and
  function names in ADVANCED_FEATURES, CONFIG_DEBUGGING,
  PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE,
  PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README
- CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use
  /config/raw/main to replace the file; describe where validation runs
- TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints
  usage)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): verify the web interface that actually ships, on port 5000

verify_installation.sh failed every healthy install: it required the
long-removed web_interface_v2.py and looked for a listener on port 5001,
while the web interface binds 5000 (web_interface/start.py). It now
checks the files ledmatrix-web.service runs (start_web_conditionally.py,
web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the
same 5001 port in its listen check, HTTP probe and printed URLs.

Port matches are anchored so :50001 no longer counts as :5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): one display-size contract: display_manager.width/height

CLAUDE.md (#580) says to read display_manager.width/height because
matrix is None when hardware init fails; the development guide, the
safety-harness doc and two DisplayManager docstrings still recommended
matrix.width/height. The bundled starlark-apps plugin read matrix.width
unguarded, so its magnify recommendation and frame scaling raised in
fallback mode (e.g. after the Pi 5 hardware refusal).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): make install_service.sh --help print usage instead of installing

install_service.sh parsed no arguments, so `sudo ./scripts/install/
install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md)
rewrote ledmatrix.service, ledmatrix-web.service and both update-verify
units and enabled/started them. It now handles -h/--help (usage, exit 0,
no changes) and rejects any other argument with exit 2 before doing
anything. Running it with no arguments, as first_time_install.sh does,
is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scroll): describe the fixed-step model and document frame_hold

Since #545 a crisp speed from scroll_config.configure() makes the helper
advance a fixed whole-pixel step per presented frame with no clock, and the
display manager's frame hold is part of the speed. The docs still described
the removed wall-clock model:

- scroll_config's module and configure() docstrings said speed is applied
  in time-based mode and that omitting the hold "falls back to fractional
  pixels"; omitting it actually runs the scroll frame_hold times too fast.
- SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both
  modes, and read a 20 ms stats median as missed refreshes although that
  is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step,
  the hold-dependent healthy median, that target_fps plays no part, and
  that a hand-added scroll_pixels_per_second loses to a schema-default pair.
- PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling)
  without frame_hold; it now documents the parameter (core 3.4.0) with a
  configure() + set_scrolling_state example.
- update_scroll_position/set_scroll_speed and set_scrolling_state
  docstrings say the same.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark target_fps legacy; describe what Vegas scroll_delay does

- General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after
  the sports_scroll fix nothing in core scrolling reads it. The field and
  its API validation stay so saved configs and plugins that read
  global_config['target_fps'] keep working. CONFIG_REFERENCE says the same.
- Vegas frame_based_scrolling/scroll_delay were described as frame-count
  stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based
  mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and
  still advances by elapsed time, so the applied speed is
  clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The
  config comments, render_pipeline comment and CONFIG_REFERENCE rows now
  say so. No behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(deps): describe how plugin dependencies are really installed

The guides said the web service runs as root, that installs pick --user
from os.geteuid(), and quoted a warning and a
PluginManager._install_plugin_dependencies() method that don't exist. The
web unit runs as the installing user; store installs go through
install_requirements_file() and sudo safe_pip_install.sh (root), with a
user-level fallback that says so, and load-time installs run in the
display service's own (root) interpreter.

Manual paths now use the configured plugins directory (plugin-repos/ by
default) instead of plugins/, which store installs no longer use, and
install_plugin_dependencies.sh is described as scanning that directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): count local changes one way for the preflight and the pull

The automatic update's preflight ignored mode-only changes and anything
whose status line contained plugins/ or plugin-repos/, then promised
"Automatic updates will not stash your changes". perform_core_update
used plain git status (modes count) and ignored only 'plugins/', then
ran 'git stash push -- :!plugins', which nothing ever pops. So an edit
to a bundled plugin under plugin-repos/, or the installer's chmods on
tracked scripts, passed the preflight and was stashed away for good.

- auto_update.local_changes() is the one predicate both use:
  core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/
  excluded by leading folder rather than substring (a core file under
  web_interface/static/v3/js/plugins/ now counts).
- Update Code's explicit stash leaves out both plugin folders; the
  pull's --autostash carries their edits and mode changes across and
  reapplies them.
- The automatic updater calls perform_core_update(stash_local_changes=
  False), which refuses instead of stashing edits that appeared after
  the preflight; update_core reports that as 'blocked'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): diagnostics follow the web autostart default and api_v3 package

#556 made a missing web_display_autostart mean "start" (only an explicit
false/off keeps the web interface down), but the diagnostics still said
otherwise: diagnose_web_ui.sh reported a missing key as "defaults to
false", diagnose_web_interface.sh said the web interface "will not start
unless this is set to true" and recommended enabling it, and
debug_web_manual.py printed False. Troubleshooting a down web UI pointed
users at a non-cause.

Both shell scripts now evaluate the setting with the launcher's own
autostart_enabled() (inline fallback if it cannot be imported) and report
on / off / not set (on) / unparseable config; debug_web_manual.py uses
the same function. They also check web_interface/blueprints/api_v3/
__init__.py: api_v3.py became a package in #553, so every healthy
checkout was reported as missing a file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(install): what install_service.sh installs; verify script port; no sudo for --help

install_service.sh installs and starts ledmatrix, ledmatrix-web and the
update-verify units, not only ledmatrix.service (systemd/README.md,
README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help'
as a harmless check; it now shows --help without sudo and warns what a
real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh
checks the web interface on port 5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note update-all, plugin system settings and script fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): size the preview after orientation and pixel mappers

display_geometry.physical_size claimed to give DisplayManager's answer but
only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured
after the library's pixel mappers, so a Rotate:90 / orientation 90 chain
previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for
128x64.

Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc
does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown
names leave it alone), and move the orientation composition here so
DisplayManager and the preview share it. The module docstring no longer
claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH.

Audit finding F18.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): refuse settings the rgbmatrix library aborts on, on every board

The library answers several settings with a NULL matrix or abort() rather
than an error, so the display service crash-looped (Restart=on-failure)
instead of reaching fallback mode: rows above 64, chain_length above 255
(uint8_t binding setter, documented as "no upper limit"), a misspelled
hardware_mapping, and parallel 2-3 on a single-output mapping, reachable
from the Display form on the default adafruit-hat(-pwm) mapping. #586 only
guarded the Pi 5 subset.

- src/matrix_support.py holds the rules for every board (Options::Validate
  ranges, binding integer types, mapping names and outputs from
  lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the
  API's numeric ranges.
- DisplayManager checks them before building options and raises
  MatrixSettingsRefused, so a hand-edited config falls back with a logged,
  reported reason. Emulator mode only warns.
- The config API refuses them with a 400 naming the setting; combinations
  are checked against stored values but reported only when the request
  sets a field involved.
- The hardware status file gains "cause" (settings/library/forced). The
  fallback log and Display banner give the Pi 5 rebuild hint only for a
  library failure instead of rebuild + gpio_slowdown advice for every
  failure; one Pi 5 slowdown recommendation (1-3, start at 1).
- The Display form offers classic/classic-pi1 and orientation 90/270 and
  renders any other stored mapping selected with a warning, so an
  unrelated save no longer rewrites them; the API accepts 90/270.

Audit findings F03, F16, F19, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(display): library limits, template defaults and Pi 5 slowdown

- rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs,
  classic/classic-pi1 mappings and orientation 90/270 documented.
- Defaults are the config.template.json values: config migration adds
  missing keys from the template, so the listed "code defaults" never
  applied.
- One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode,
  starting at 1.
- Troubleshooting describes the refused-settings fallback, and CHANGELOG
  corrects the Unreleased "no upper limit" entry.

Audit findings F19, F20, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py opens the panel with the service's options

--measure and --demo built RGBMatrixOptions from a private copy of the
display service's builder that had drifted: gpio_slowdown came from
display.hardware (default 2) instead of display.runtime (default 3), and
rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors,
pixel_mapper_config and orientation were skipped, with different defaults
(hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown
was measured -- or garbled -- in a setup the service never drives.

The option filling in DisplayManager._setup_matrix moves, unchanged, into
DisplayManager.apply_matrix_options(options, config), which _setup_matrix
calls and the script reuses (overriding only limit_refresh_rate_hz for
--measure). The script now loads the whole config rather than the hardware
block. Tests pin the script's options to the service's attribute for
attribute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py recommends keys the resolver honours

The ladder ended by telling users to set
display_options.scroll_pixels_per_second. scroll_config ranks that key
below the scroll_speed + scroll_delay pair, deliberately, and several
plugin schemas default the pair into config, so the advised key was
silently ignored (a schema-default 1/0.02 pair plus an advised 66 still
resolved to 50 px/s).

The advice is now the pair that selects the crisp speed exactly
(pixels_per_frame every frame_hold/refresh seconds), explains that the
pair outranks scroll_pixels_per_second, and gives the scoreboards'
per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the
printed pair over a schema-default pair and check it lands on the
advertised speed and hold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula

- SPORTS_UNIFICATION.md still presented honouring global target_fps as
  sports_scroll's added behaviour and its one user-visible gain; note that
  it was withdrawn because it had become a speed multiplier.
- ADVANCED_FEATURES.md gave Vegas scrolling as
  (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s
  by elapsed time, through a 0.1-5 px per scroll_delay clamp when
  frame_based_scrolling is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): scroll model fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(dev): link-github links plugins from the ledmatrix-plugins monorepo

link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git,
and those per-plugin repositories no longer exist: official plugins are
directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the
monorepo once into the dev directory, finds plugins/<name>,
plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and
links it under its manifest id. With an explicit repo URL it still links a
single-repository plugin as before.

dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a
fork), plus plugins_repo and plugins_branch; github_pattern, which was
documented but never read, is dropped and warned about. Ships
dev_plugins.json.example and git-ignores dev_plugins.json, both of which
the guide promised. Reading JSON falls back to python3 when jq is missing
(get_plugin_id silently returned nothing without jq).

update/status/list find the git checkout above a monorepo plugin
directory (its .git is not in the plugin dir), and update pulls a shared
checkout once. status no longer exits 1 when nothing is broken.

Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration,
workflow, store integration, hello-world link, submission), and the
nonexistent scripts/git-hooks/pre-push-plugin-version and
scripts/bump_plugin_version.py replaced with the real rule: bump the
manifest version and run update_registry.py. scripts/dev/README.md and
CLAUDE.md updated to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scripts): monorepo workspace layout; fix_perms and install READMEs

MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin;
setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into
plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the
workspace file opens LEDMatrix plus ../ledmatrix-plugins.

scripts/fix_perms/README.md listed cache directories
fix_cache_permissions.sh never touches and a 'ledmatrix' service user
that doesn't exist (also in scripts/install/README.md); adds
safe_pip_install.sh. install/README: install_service.sh installs the web
and update-verify units too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): keep the rollback's pip retries inside the unit time limit

The health check reinstalled the previous requirements by trying the
next bash path after any failure, including a 600 s pip timeout. Two
files, two paths: up to 40 minutes of pip alone, while systemd stops
ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the
rollback half-way and leaving the update 'verifying' until the web UI
calls it lost.

- Like permission_utils.install_requirements_file, only a sudo refusal
  moves on to the next bash; a pip that ran and failed or timed out is
  not repeated. The refusal wording is one list
  (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only
  verifier and pinned equal by a test.
- All reinstalls in one rollback share a 600 s budget.
- WORST_CASE_SECONDS adds up every timeout on the longest path (27.5
  min); a test holds it under the unit's TimeoutStartSec and that under
  the web UI's VERIFY_LOST_SECONDS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools

Plugin config was prepared differently depending on how it arrived:

- JSON POST /plugins/config built a partial body on schema defaults, so
  {"enabled": true} reset every other setting of the plugin. It now merges
  onto the stored section first, as the form path already did.
- Legacy-boolean normalization (#588) ran only at load: GET /plugins/config
  returned the raw boolean, posting it back failed validation, and hot
  reload handed plugins the raw section (a legacy dynamic_duration: true
  came back as a boolean). schema_manager.prepare_plugin_config (normalize,
  then defaults) is now used by PluginManager.load_plugin, both save paths,
  GET, the save notifications and DisplayController's hot-reload callback.
- The JSON save's filter kept only enabled/display_duration/live_priority
  and dropped a submitted skin, skin_options or vegas_* tuning key. There
  is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES,
  used by validation and by the save filter; PluginManager's
  CORE_OWNED_CONFIG_KEYS is its vegas subset.
- Plugin sections posted to /config/main were stored verbatim, including
  values /plugins/config rejects. They now go through the same preparation
  (_prepare_plugin_config_for_save, extracted from save_plugin_config), and
  a failing section rejects the whole save before anything is written.
- dev_server read only top-level defaults and let a schema enabled:false
  win; build_full_config shallow-merged overrides, dropping sibling
  defaults; the harness extracted defaults differently from the device.
  loading.build_config now uses the device's extraction and preparation,
  and dev_server, check_plugin, render_plugin and the harness all use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt-bridge): brightness changes apply live and touch nothing else

The display service's hot reload applies a saved brightness within a few
seconds, and /config/main no longer resets other display settings on a
brightness-only JSON body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): automatic update hardening

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI

It described web_interface_v2.py and index_v2.html (both gone), client-side
form generation, one POST per field with {key, value}, and 'no nested
objects'. The v3 UI renders plugin forms server-side from the schema
(pages_v3 partial + plugin_config.html macros, nested sections and
x-widgets), posts the whole form once, and save_plugin_config() merges onto
the stored section, validates, splits x-secret fields and notifies the
plugin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt): brightness saves apply via hot reload and leave other settings alone

The bridge README said brightness is applied on the display's next
restart; the display controller's config hot reload applies it within
seconds. It also now states that the bridge's partial JSON save changes
only brightness (the /config/main merge fix in this change set).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): don't log pip's output from the health check's reinstall

pip can echo a private index URL with embedded credentials;
permission_utils redacts it, the stdlib-only verifier cannot, so it
logs the exit code only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark the plugin_system toggles as unused legacy keys

auto_discover, auto_load_enabled and development_mode are read by
nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE
said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and
the REST reference listed them as live settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): docs and developer tools group

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): legacy plugin-system toggles no longer count as a General save

auto_discover, auto_load_enabled and development_mode have left the General
form, so a post carrying only one of them is not a general-settings save and
must not treat web_display_autostart and auto_update as unchecked. The
plugin_system block itself is left as on main for the branch that reworks it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): config-save and plugin-config preparation fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(claude): re-check matrix_support.py rules when the library submodule is bumped

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address Codacy findings on the core audit PR

- plugin_manager.prepare_plugin_config: when the fallback legacy-boolean
  pass also fails, log a warning instead of a bare except/pass.
- api_client.js: request() refuses any endpoint that is not a plain path
  under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace,
  control characters) with INVALID_ENDPOINT before calling fetch(), and
  plugin ids are URL-encoded wherever they are put into a URL (also in the
  app-shell batch load).
- test_update_all.js: pins both against the shipped client.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): check endpoint control characters without a control-character regex

Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags
the \x00-\x1f range in checkEndpoint's regex. Test the char codes
instead; the endpoints refused are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(auto-update): make the seed script executable on disk, not only in the index

On Linux Repo.publish() commits with -a, which recorded scripts/run.sh
as 100644 upstream because the seed file was never chmod +x. The pull
then brought in the same mode the installer chmod had made locally, so
installer_chmod saw no mode change left to check. The updater was fine:
with the upstream commit at 100755 the --autostash carries the device's
chmod across. Verified under Linux (WSL, git 2.43): the old helper fails
exactly as CI did, the fixed one passes all 63 tests in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 16:37:29 -04:00
ant456 a3d505384d Add render_width/render_height support to Starlark Apps (#552) 2026-09-11 11:22:52 -04:00
ChuckandClaude Opus 5 bdb9a94033 refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package

web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181
functions in one module -- 9% of the core by line count and three times the
next largest file. It becomes a package of nine route modules grouped by path
segment, plus __init__.py for the shared imports, constants, Blueprint and
helpers.

Every route module decorates the SAME api_v3 Blueprint object, so endpoint
names stay api_v3.<function>, the URL map is unchanged and app.py is untouched.
Verified: 111 routes before, 111 after, byte-identical rules, endpoints and
methods, and every endpoint still on the one blueprint.

  plugins   3,867   config    1,178   starlark  692   system  619
  fonts       452   misc        398   wifi      361   display 326   backup 212
  __init__  1,787 (imports, constants, Blueprint, 56 helpers)

Two things the URL-map check could not catch, both found by running the suite:

1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one
   directory deeper made that resolve to web_interface/ instead of the project
   root. Nothing failed at import; it surfaced as ~110 tests failing with 404s
   and "installation script not found", because every path built from it was
   one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts
   PROJECT_ROOT/run.py exists so the next move cannot repeat it.

2. Module-attribute patching. Tests do
   monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route
   module that binds such a name by value never sees the patch. The shared code
   therefore stays in __init__.py rather than moving to a _common submodule --
   it has to live on the module the tests patch -- and the eleven names tests
   patch are read back through the package (_pkg.X) instead of bound by value.
   Those eleven were found by AST-scanning every setattr in the test tree, not
   by guessing; "time" is among them, used to drive a fake clock through the
   second-resolution credential-backup filenames.

Test changes are confined to what genuinely moved: patch targets that now name
the owning route module, imports of helpers, and six tests that scan the api_v3
source as a file and now read the package directory.

Full suite: 4,278 passed, 68 skipped, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(api-v3): address CodeRabbit findings from the blueprint-split review

Fixes to the api_v3 package split (PR #553), one per finding verified
against the actual code:

- __init__.py: _redact_credentials only blanked scalar values under a
  credential-named key; a bare list of secrets under such a key (e.g.
  tokens: ["a", "b"]) passed through untouched, since the list branch
  recursed with no memory that its key looked like a credential. Nested
  dicts still walk normally (a documented, tested behaviour -- a container
  like secrets: {api_key: ..., note: ...} is a section name, not a value to
  blank outright), but any value reached under a credential-shaped key is
  now actually blanked.

- __init__.py: the OAuth helper script's raw stderr/stdout went to
  logger.error unredacted (CWE-532) right next to a comment claiming this
  was deliberate; the HTTP response already used the existing redact_text
  helper. Routed the log line through the same helper.

- __init__.py / starlark.py: the standalone Starlark manifest fallback
  (used when the plugin instance isn't loaded) read-modified-wrote
  manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe
  (plugin-repos/starlark-apps/manager.py), which already holds an flock for
  the same file when the plugin is loaded. Added _starlark_manifest_lock,
  mirroring that pattern, and wrapped every standalone read-modify-write
  call site in it. The app-config update route also wrote config.json and
  the manifest as two separate, non-transactional writes (a second,
  distinct finding at the same call site); config.json is now rolled back
  if the manifest write that follows it fails.

- backup.py: restore options used bare bool() on values from the request,
  so {"restore_secrets": "false"} restored secrets anyway (bool("false") is
  True). Switched to the existing _coerce_to_bool helper already used for
  this exact purpose elsewhere in the package.

- config.py: an automated import-rewrite mangled four user-facing
  validation strings and their neighbouring comments -- "Invalid start
  time" had become "Invalid start _pkg.time" (and likewise for "end time")
  in both the schedule and dim-schedule per-day validation paths.

- display.py: `import _pkg.time as time_module` -- _pkg is a local alias
  for the package, not a real importable module, so this raised
  ModuleNotFoundError whenever a caller restarted an already-running
  display service via /display/on-demand/start, after the on-demand
  request was already written to cache. Fixed to `import time`. Audited
  the rest of the package for the same `_pkg.<module>` import mistake;
  every other `_pkg.` reference is a legitimate attribute read-through
  (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken
  import statement.

- fonts.py: validate_file_upload's max_size_mb parameter is silently
  unused by that helper (it only checks filename/extension) -- the font
  upload route saved arbitrarily large files as a result. Added the same
  seek-and-check pattern already used for the sibling .star upload.

- wifi.py: two ad hoc, inconsistent bool coercions. POST
  /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled
  it. POST /wifi/radio's enabled/force parsing recognized real bool and
  some strings but not int 1/0 (1 is True is False in Python). Factored one
  small _parse_bool_ish helper local to this file and used it at all three
  sites.

Not changed: the "unknown/misspelled restore option keys default to True"
half of the backup.py finding -- the file's own comment documents that a
missing key deliberately means "restore everything," matching the
already-existing JSON-parse-failure guard a few lines above it; only the
bool-coercion defect was a real bug.

Added or extended regression tests for every fix, following each area's
existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed
on both this branch and origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to this
change) -- no new failures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* fix(api-v3): reject unknown restore option keys

CodeRabbit's review of the blueprint split (#553) asked that
POST /backup/restore reject option keys outside RestoreOptions'
known set. The follow-up commit fixed the bool("false")-is-True
bug with _coerce_to_bool but never added the key check: a typo'd
or renamed key (e.g. "restoreSecrets") is silently ignored by
opts_dict.get(key, True), so the flag stays at its True default
and secrets get restored despite the caller's request saying
otherwise -- with no indication anything was wrong.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb

* fix(api-v3): address CodeRabbit findings on the blueprint split

- _redact_credentials: blank scalar descendants of objects reached
  through a credential-owned list (e.g. tokens: [{"value": "secret"}])
  regardless of field name -- the existing name-based walk only
  protected direct dict values under a credential key, not list items.
- wifi.py: reject enabled/force/auto_enable_ap_mode values
  _parse_bool_ish can't recognize (400) instead of silently treating
  them as False, which could disable Wi-Fi or the radio itself.
- Starlark manifest locking: lock a stable manifest.json.lock sidecar
  instead of manifest.json itself, in both the standalone route path
  (_starlark_manifest_lock) and the plugin path
  (StarlarkAppsPlugin._save_manifest / _update_manifest_safe).
  manifest.json is replaced by an atomic rename on every write, which
  swaps in a fresh inode; a lock held on the old inode does not
  exclude a second locker that opens the path afresh right after the
  rename and gets the new inode, so two writers could race despite
  each holding "a lock". A sidecar that no write ever touches always
  resolves to the same inode for every locker.

Skipped as stale: the "serialize the complete manifest
read-modify-write" finding at api_v3/__init__.py -- every standalone
handler that calls _write_starlark_manifest is already wrapped in
_starlark_manifest_lock() on this branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): re-check reconciliation findings by the reconciler's own rules

Both CodeRabbit findings on the merge commit, verified against the code first.

Major, plugins.py: the stale-findings filter derived its own notion of "in
config" and "on disk", and both were looser than the reconciliation module's.
set(load_config()) also contains system keys, the secrets-file keys load_config()
merges in, and non-dict values; and any directory holding a manifest.json
counted as installed even when that manifest does not parse. Either looseness
clears a finding that is still true -- and a secrets key read as a plugin is the
precise bug the filter exists to stop reporting, so reintroducing that asymmetry
while re-checking was the wrong way round.

The two extractions now live in state_reconciliation.py as config_plugin_ids()
and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys()
alongside. _get_config_state() and _get_disk_state() use them too, so there is
one definition rather than two that can drift. _get_disk_state() re-reads each
manifest for version/name after taking membership from the shared extractor;
that costs one extra small read per plugin on a path that runs once per boot.

Minor, the new test: the fixture assigned api_v3.config_manager and
api_v3.plugin_manager directly. Those live on a module-level blueprint
singleton, so the mocks leaked into every later test that imports api_v3 --
pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr,
which restores them. This is the same pollution class that made an earlier test
in this session break seven unrelated ones, so it is worth getting right.

Five cases added for the parity itself: a secrets key, a system key and a
non-dict value must not clear an "installed but missing from config" finding,
and neither an unparseable manifest nor a .standalone-backup- directory may
count as installed. All five fail against the looser version.

Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and
CodeRabbit all pass. Codacy reads action_required on every commit of this
branch including the first, so it is pre-existing and not from this work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 10:07:38 -04:00
ChuckandClaude Opus 5 e23f1f45d3 feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge

Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of
patches and services built while running this project on Starlark apps
under MQTT control. Its patches are whole-file copies taken against an
older tree, so applying them as written would revert #523's frame
pacing, #534's display() bool returns and the GitHub token masking in
plugins_manager.js. Three of its claimed fixes are already in main, and
its api_v3 Starlark routes are #535's. What follows is the rest --
verified against current code, and reimplemented where the patch's
approach did not hold up.

**On-demand display.** `pinned` reached the controller from the API, was
stored on it and republished in the status payload, but never narrowed
the rotation -- a pinned request still cycled every mode its plugin
owns. Right for a sports plugin, whose modes are views of one subject;
wrong for a plugin whose modes are unrelated, which is every Starlark
app. Now honoured, and it survives a restart.

Restarting while on-demand was active loaded *only* the on-demand
plugin, so normal rotation had nothing to return to for the life of the
process -- and a restart mid-session is routine, since that is how an
update is applied. The panel came back cycling one plugin's modes with
no way out but clearing the cache by hand. Every enabled plugin loads
now; on-demand still resumes on its saved mode.

Stop requests are exempt from the duplicate guards on purpose, so that a
second click stops a mode a race left running -- which means consuming
the mailbox is the only thing that ends one. It was never consumed, so
the same stop was re-read and re-processed on every poll, forever. Both
paths now share one compare-before-delete helper.

**Starlark rendering.** `extract_schema` parsed the source with a regex,
which can only see option lists written out literally: an app whose
dropdown is filled from a live API call inside `get_schema()` came back
empty, and the config form offered nothing to pick. Now runs `pixlet
schema`, which executes the app, and falls back to the parser when
Pixlet is absent, too old for the subcommand, or the app fails to run.
The third-party patch replaced the parser outright and hardcoded
/usr/local/bin/pixlet; this keeps the fallback and the binary search.

A `|` in a config value was dropped by a shell-metacharacter filter,
though the command is a list with no shell involved -- and apps do use
it as a separator inside one value. The key went missing silently and
the app rendered its own "not configured" screen with nothing to say
why. And a 0-byte render was reported as success: Pixlet exits 0 and
writes nothing when an app has no content, which read downstream as a
working app drawing a black panel.

**Starlark display.** `display()` ignored the mode it was called with,
so a specific app could not be addressed. It now accepts `display_mode`
-- which is the whole mechanism, since the controller inspects the
signature before passing it. Found while there: `_select_next_app` ran
only while `current_app` was unset, so with several apps installed the
first was picked once and shown forever while the rest were rendered on
schedule and never displayed. And `enable_scrolling` was missing, so
multi-frame apps were called once per rotation slot and never advanced
past frame one.

**GET /api/v3/display/modes.** Every mode that can be requested
on-demand, with the plugin that owns it. Nothing exposed this, so
anything driving the display from outside the web UI read each plugin's
manifest.json off disk and reimplemented PluginManager's fallbacks. It
also triggers discovery, which is otherwise lazy and normally happens
because a person opened the dashboard.

**integrations/mqtt_bridge.** Home Assistant control over MQTT
Discovery: a mode select, a stop button, power, brightness. Rewritten
against the API rather than the filesystem, so it needs no read access
to config.json and cannot drift from the web UI. paho-mqtt 2.x
VERSION2, TLS, an availability topic that is also the last will, and
secrets from the environment.

**Two opt-in extras.** A DNS single-request unit, for glibc's parallel
A/AAAA lookup stalling ~5s per name on routers that answer only the A
query -- which makes any plugin calling an external API slow and
Starlark apps, which have a render timeout, fail outright. And a Pixlet
config editor: a script you run and Ctrl+C rather than the third-party
version's always-on unauthenticated Flask service, since it stops the
display for the length of a session. Neither is installed by default.

Long Starlark app names now wrap instead of overflowing their card.

115 new tests across 5 files. Also unblocked
test_starlark_display_contract.py, which was silently skipping wherever
fcntl is absent. Whole suite: no new failures against main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(mqtt_bridge): the five issues Codacy flagged on this branch

All in the new bridge, all real:

  * requests floor was 2.31.0, which carries CVE-2024-35195,
    CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which
    is what the project's own requirements.txt already pins.
  * `import time` was never used.
  * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential.
    It is the "no password configured" default; marked nosec B105, the
    convention used elsewhere in the repo.

Also dropped an unused `build_app` from the display-modes test imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: the review findings on this PR

Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and
is answered below.

**One bad config section blanked the whole mode list.**
`/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`,
so a non-dict under a plugin id -- a shape DisplayController already
guards, so it happens -- raised AttributeError mid-loop and answered 500
with no modes at all. Every MQTT bridge entity is built from that list.
Now skipped with a warning.

**The DNS scripts reported success they had not earned.** Three separate
paths: `resolvconf -u` failing was swallowed by `|| true`; the
systemd-resolved branch exited 0 without applying anything, so the
oneshot unit recorded success while the workaround was inactive; and the
installer's `|| echo` turned a failed start into "installation
complete." with exit 0. All three now fail loudly. `single-request` is a
glibc resolv.conf option with no resolved.conf equivalent, so on those
hosts the honest answer is that it cannot be applied.

A NetworkManager-generated resolv.conf is regenerated on connection
changes, not only at boot, and the unit is oneshot with RemainAfterExit
-- so the option can vanish mid-boot with nothing to put it back. Now
detected and stated plainly rather than implied to be permanent.

**`Before=` does not order a manual restart.** It only orders units
already in the same transaction, so `systemctl restart ledmatrix` could
bypass the fix. install_dns_fix.sh now writes a ledmatrix.service
drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround
failing should not stop the display.

**The Pixlet editor's `--lan` is gone.** `pixlet serve` has no
authentication, and a printed warning is not access control. Loopback
only, with the SSH port-forward in the header where the flag used to be
documented -- SSH does the authenticating and nothing is left listening.

**The MQTT example config now defaults to TLS** on 8883. The installer
copies it verbatim, and without TLS the broker password and every
command cross the network in cleartext. A plaintext broker is still
supported and documented, and the bridge warns once at startup when a
password is configured without TLS.

**Not taken: "the upstream Pixlet CLI has no `schema` subcommand."**
Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh`
installs `tronbyt/pixlet`, whose `cmd/schema.go` is
`schema [PATH]` -> JSON on stdout, built on
`runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is
exactly what extract_schema_via_pixlet calls. A binary without the
subcommand exits non-zero and falls back to the source parser, which is
already covered by a test.

**CodeQL stack-trace exposure: not taken either.** I removed `details`
first and that broke
test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception,
which enforces `describe_exception` across all ~75 handlers -- written
because a device with failing storage answered "see logs for details"
from the log viewer itself. describe_exception redacts credentials; the
trade-off is the project's and is already made. Restored, with the
reasoning in a comment.

11 new tests. Whole suite: no new failures against main, 4127 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 08:34:52 -04:00
ChuckandClaude Opus 5 a0d3e64099 fix: ten defects found validating the whole plugin fleet on hardware (#534)
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers

Three independent fixes found while validating every plugin on a 256x64 rig.

FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf
ships with the core, so every plugin offering it logged "Font family
'tom_thumb' not found" (16 warnings per countdown render) and had to carry a
private loader to use a bundled font. Closes #524.

VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter
that DisplayManager gained, so any plugin passing it died with TypeError at
render time and failed every size. Nine plugins now make that call;
ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other
seven only passed because their scroll path was unreachable without data.
Closes #525.

api_v3 declared module-level config_manager/plugin_manager = None that nothing
ever assigned -- app.py sets the blueprint attributes, which the other 150+
call sites use. Three sites read the decoys, so /health reported the config
unreadable and the plugin system uninitialised (making "degraded" permanent and
unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on
every rig. The decoys are removed rather than assigned, so a bare name is now a
NameError at test time instead of a silent None. The same function's first-call
uptime was computed from two separate clock reads and came out negative.
Closes #529.

Verified on the rig: both previously-failing plugins render, the tom_thumb
warnings are gone, /health reports "healthy" with all three checks passing, and
/display/current reports the real 256x64.

* fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark

The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is
world-writable and sticky, and the display service runs as a different user
from the tooling, so a leftover temp owned by anyone else became unopenable
even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a
sticky directory. The preview and the health check's liveness proxy then froze
until someone deleted the file by hand; on the test rig that meant 23 hours of
a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with
cleanup on failure, matching the hardware-status write a few hundred lines
above. Closes #528.

_poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults
the in-memory TTL to max_age -- so the first request was pinned in memory for an
hour and every later poll returned that stale copy. No second on-demand request
was honoured until the service restarted, while the API kept returning 200.
get() already documents memory_ttl=0 for exactly this cross-process case.
The consumed request is also now deleted: leaving it on disk meant a restart
replayed the previous request, activated it, and ignored the one the caller had
just made. Closes #530.

starlark-apps returned None from display() when it has no app to show, which is
the state of every install without Pixlet and of a fresh one before any app is
added. The controller only skips on a boolean False, so that held a black panel
for the full display_duration instead of rotating on. Closes #456 (core side).

Verified on the rig: two consecutive on-demand requests with no restart between
them are both activated, where the second was previously dropped in silence.

* perf(harness): share one cache across a plugin's renders

_instantiate built a fresh MockCacheManager for every (size, mode), and that
mock is a per-instance in-memory dict, so each render was a cold start. A plugin
that fetches per game or per player re-fetched everything N times over --
baseball-scoreboard at one size took 840s for nine renders where the arithmetic
said ~72s, and at eight sizes it exceeded a 900s timeout.

The second and later renders also never exercised the cache-hit path, which is
what a running rig executes almost all of the time, so a caching regression
could not be caught here.

The cache is now built once per render_plugin_matrix call and threaded down.
The display manager stays per-render -- the bounds checking depends on that --
so only fetched data is shared.

Measured on the rig, same render counts and same goldens:
  tide-display         2s -> 1s   (32 renders)
  cricket-scoreboard  10s -> 3s   (24 renders)
No pass/fail change across tide-display, cricket-scoreboard, clock-simple,
geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages.

Closes #533.

* fix(scripts): run standalone plugin tests instead of collecting nothing

run_plugin_tests.py discovered every plugin test file and handed the lot to
pytest. Most plugin tests are standalone scripts -- module-level main() plus an
`if __name__ == "__main__"` guard, signalling through an exit code -- and pytest
collects zero items from those. The run printed how many files it had *found*,
then "no tests ran", and exited without executing any of them. On a rig with all
44 first-party plugins that is 151 of 248 files.

Files are now classified and each kind runs under the right runner: pytest for
real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1
fail convention ledmatrix-plugins' own runner established (a script that wants a
tty or an LED matrix is a skip, not a regression).

Before:
    $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos
    Found 1 test file(s)
    collected 0 items
    no tests ran in 0.31s                      rc=0

After:
    Found 1 test file(s) -- 0 collectable, 1 standalone script(s)
    1 passed, 0 skipped, 0 failed (scripts)    rc=0

Verified across three shapes: countdown (1 script), jellyfin-now-playing and
pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files
split 4 collectable / 7 scripts, all seven of which had never run).

Closes #532.

Running the flights scripts for the first time also surfaced four genuinely
failing tests there, hidden by the mirror-image bug in the plugins repo's own
runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465.

* fix(harness): give an empty-looking mode a few frames before warning about it

check_plugin's "drew nothing but display() returned X" warning fired on a single
frame, rendered with force_clear=True, under a frozen clock. All three defeat a
scrolling plugin, whose first frame is legitimately its blank scroll-in buffer.
Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which
people stop reading a warning, which matters because the true positives are
real: a mode that draws nothing and does not return False holds a blank panel
for its whole display duration.

An apparently-empty frame is now re-driven for up to 48 more frames with
force_clear=False (force_clear means "reset the scroll", so repeating it would
redraw frame 1 for ever) and with the clock advancing -- freezegun's factory
where time is frozen, a real sleep where it is not, since scroll position is
usually a function of elapsed time. The first frame that draws content replaces
the result.

The clock is moved back afterwards. It is shared by every render in the matrix,
so time borrowed by the probe leaked into later modes and drifted their goldens
-- f1_upcoming picked up 5 spurious drifts before this was restored.

Measured on the rig:

                        empty warns          check
                        before  after
  f1-scoreboard            42      0    48 PASS / 0 FAIL, goldens intact
  ledmatrix-elections      16      0    16 PASS / 0 FAIL
  on-air                    8      8    true positive, kept
  nfl-draft                 8      8    true positive, kept
  clock-simple/geochron/    0      0    unchanged
  christmas-countdown

58 false positives gone, both true positives kept, no golden regressions. Cost
is confined to modes that really are blank: plugins that draw immediately are
unchanged (clock-simple and tide-display still 2s), while on-air -- eight
deliberately blank modes -- goes to 21s.

Closes #527.

* fix(harness): load nested schema defaults, and merge caller config at leaf level

load_config_defaults read only top-level properties. An object property carries
its defaults on its children, not on itself, so everything nested was dropped --
2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of
565. render_plugin_matrix's comment says the plugin then "behaves like a real
install", which for most of the fleet it did not.

_defaults_from_properties now recurses. merge_config deep-merges the caller's
config onto the result so an override lands at the leaf: a shallow merge would
let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard
every other nhl default, which is the same class of bug being fixed here.

Measured before/after across all 49 installed plugins on the rig: **no render
changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact.
Plugins already fall back to the same values internally via config.get(key,
default), so supplying them explicitly agrees with what they were doing. The
defaults really are arriving now:

  ufc-scoreboard        9 -> 87 defaults
  ledmatrix-flights    51 -> 95
  masters-tournament   10 -> 51
  cricket-scoreboard   22 -> 50
  tide-display         12 -> 18

and hockey-scoreboard, which used to load nhl.enabled=None, now gets
nhl.enabled=True with its full display_modes block.

Caveat worth carrying: the eight plugins with the most nested config
(soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of
the 2,386 dropped defaults, 68%) could not be measured. They import
src.common.sports_shared, which the test rig's core branch predates, so they
fail to load there identically before and after. Re-run this comparison against
a core that has that module before trusting the "nothing changed" result for
them; those are exactly the plugins whose renders should change most.

Closes #531.

* refactor: narrow the exception handlers this branch introduced

Codacy flagged the new code; it passes on other recent PRs, so the finding is
mine. Four of the five broad `except Exception` clauses I added were catching
far more than they needed to, which is the same shape as several bugs this
branch fixes -- hello-world's TypeError sat invisible for exactly this reason.

  freezer() / move_to() / tick()   -> (AttributeError, TypeError, ValueError)
  cache_manager.delete()           -> (OSError, AttributeError, KeyError)

The fifth stays broad and now says why: it wraps a call into a plugin's own
display(), which can raise anything, and the first frame has already rendered --
so a failure there must not turn a good result into an error.

Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty
warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows
7 golden drifts both before and after this branch, so it is not from these
changes -- its committed goldens predate #521's 1-bit text rendering.

* fix: resolve CodeRabbit review and Codacy findings on #534

CodeRabbit raised six; all six were real.

The test double had drifted ahead of production. VisualTestDisplayManager
accepted set_scrolling_state(frame_hold=...) while DisplayManager did not,
so such a call passed every harness run and would raise TypeError on the
panel -- the one failure a safety harness exists to prevent. frame_hold
belongs to the change that adds it to DisplayManager (#523), so it moves
there and the double matches main again.

The harness swallowed exceptions from re-rendered frames. _settle_loop
re-renders a mode that came back blank, to give a scroll time to draw;
returning silently on a crash meant a mode that renders one good frame
and then explodes was reported as passing. Recorded on result.error now,
keeping the captured frame so the failure stays inspectable.

starlark-apps display() returned True after _display_frame() failed, so
the controller held a dead frame for the whole display_duration instead
of rotating on. _display_frame now returns bool on all three paths.

run_plugin_tests.py used env.setdefault for PYTHONPATH and
LEDMATRIX_CORE, so an inherited value won and the subprocess imported a
different core than the one under test -- ledmatrix-plugins#467 exactly.
Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally.

The on-demand mailbox is polled after every frame, ~125x/second on a
scrolling mode, and the read is deliberately uncached, so it was that
many disk reads per second to find nothing. Floored at 250ms, which is
imperceptible for a web-UI click. Consuming it also deleted whatever was
present rather than what had just been processed, so a request posted
while the previous one was in flight was thrown away and never ran; the
delete is now keyed by request_id. That narrows the window rather than
closing it -- a true atomic claim needs a primitive the cache layer does
not offer, and the code says so rather than implying otherwise.

Codacy's 2 criticals were bandit B404/B603 on the subprocess call added
to run_plugin_tests.py. Fixed interpreter, argument list, no shell;
annotated with the repo's existing nosec convention. Bandit is clean on
the file.

Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py
(4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of
those fail against the pre-fix code. Full suite: 3961 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: satisfy Codacy's subprocess checks on the new test runner

Codacy runs Bandit and Opengrep (its Semgrep fork). The new
subprocess.run in scripts/run_plugin_tests.py trips three patterns, on
two different lines:

  Bandit   B404 on the import, B603 on the call
  Opengrep dangerous-subprocess-use-audit          on the run( line
           dangerous-subprocess-use-tainted-env-args on the argv line

A nosemgrep applies only to its own line, so the call line and the argv
line each need one; a single comment on the call covered neither rule
fully. Suppression is the right answer here rather than a rewrite: the
interpreter is sys.executable, the arguments are a list, and no shell is
involved, so there is nothing to word-split or expand.

Matches the pair the rest of the repo already uses for this shape --
permission_utils.py, plugin_loader.py, install_dependencies_apt.py.

Codacy: 0 new issues, up to standards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: leave visual_display_manager untouched so #523 can merge

The only change this branch made to that file was a docstring, and it
collided with #523's rewrite of the same method -- so #534 and #523 each
merged cleanly against main but conflicted with each other. Reverted to
main's text; #523 owns this method and adds frame_hold to it.

The note the docstring carried ('frame_hold arrives in #523') would have
been stale the moment #523 landed anyway. The parity test in #523 is
what actually keeps the two signatures honest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: add the Ruff suppression nosec/nosemgrep do not cover

Ruff reports S603 on the same call Bandit and Opengrep do, and none of
the three suppressions covers the others. Confirmed the precondition
first: path comes from discover_plugin_tests(), which globs test files
inside the repo, and the call is a fixed interpreter with a list argv
and no shell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:37:38 -04:00
302235a357 feat: Starlark Apps Integration with Schema-Driven Config + Security Hardening (#253)
* feat: integrate Starlark/Tronbyte app support into plugin system

Add starlark-apps plugin that renders Tidbyt/Tronbyte .star apps via
Pixlet binary and integrates them into the existing Plugin Manager UI
as virtual plugins. Includes vegas scroll support, Tronbyte repository
browsing, and per-app configuration.

- Extract working starlark plugin code from starlark branch onto fresh main
- Fix plugin conventions (get_logger, VegasDisplayMode, BasePlugin)
- Add 13 starlark API endpoints to api_v3.py (CRUD, browse, install, render)
- Virtual plugin entries (starlark:<app_id>) in installed plugins list
- Starlark-aware toggle and config routing in pages_v3.py
- Tronbyte repository browser section in Plugin Store UI
- Pixlet binary download script (scripts/download_pixlet.sh)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(starlark): use bare imports instead of relative imports

Plugin loader uses spec_from_file_location without package context,
so relative imports (.pixlet_renderer) fail. Use bare imports like
all other plugins do.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(starlark): make API endpoints work standalone in web service

The web service runs as a separate process with display_manager=None,
so plugins aren't instantiated. Refactor starlark API endpoints to
read/write the manifest file directly when the plugin isn't loaded,
enabling full CRUD operations from the web UI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(starlark): make config partial work standalone in web service

Read starlark app data from manifest file directly when the plugin
isn't loaded, matching the api_v3.py standalone pattern.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(starlark): always show editable timing settings in config panel

Render interval and display duration are now always editable in the
starlark app config panel, not just shown as read-only status text.
App-specific settings from schema still appear below when present.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(store): add sort, filter, search, and pagination to Plugin Store and Starlark Apps

Plugin Store:
- Live search with 300ms debounce (replaces Search button)
- Sort dropdown: A→Z, Z→A, Category, Author, Newest
- Installed toggle filter (All / Installed / Not Installed)
- Per-page selector (12/24/48) with pagination controls
- "Installed" badge and "Reinstall" button on already-installed plugins
- Active filter count badge + clear filters button

Starlark Apps:
- Parallel bulk manifest fetching via ThreadPoolExecutor (20 workers)
- Server-side 2-hour cache for all 500+ Tronbyte app manifests
- Auto-loads all apps when section expands (no Browse button)
- Live search, sort (A→Z, Z→A, Category, Author), author dropdown
- Installed toggle filter, per-page selector (24/48/96), pagination
- "Installed" badge on cards, "Reinstall" button variant

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(store): move storeFilterState to global scope to fix scoping bug

storeFilterState, pluginStoreCache, and related variables were declared
inside an IIFE but referenced by top-level functions, causing
ReferenceError that broke all plugin loading.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(starlark): schema-driven config forms + critical security fixes

## Schema-Driven Config UI
- Render type-appropriate form inputs from schema.json (text, dropdown, toggle, color, datetime, location)
- Pre-populate config.json with schema defaults on install
- Auto-merge schema defaults when loading existing apps (handles schema updates)
- Location fields: 3-part mini-form (lat/lng/timezone) assembles into JSON
- Toggle fields: support both boolean and string "true"/"false" values
- Unsupported field types (oauth2, photo_select) show warning banners
- Fallback to raw key/value inputs for apps without schema

## Critical Security Fixes (P0)
- **Path Traversal**: Verify path safety BEFORE mkdir to prevent TOCTOU
- **Race Conditions**: Add file locking (fcntl) + atomic writes to manifest operations
- **Command Injection**: Validate config keys/values with regex before passing to Pixlet subprocess

## Major Logic Fixes (P1)
- **Config/Manifest Separation**: Store timing keys (render_interval, display_duration) ONLY in manifest
- **Location Validation**: Validate lat [-90,90] and lng [-180,180] ranges, reject malformed JSON
- **Schema Defaults Merge**: Auto-apply new schema defaults to existing app configs on load
- **Config Key Validation**: Enforce alphanumeric+underscore format, prevent prototype pollution

## Files Changed
- web_interface/templates/v3/partials/starlark_config.html — schema-driven form rendering
- plugin-repos/starlark-apps/manager.py — file locking, path safety, config validation, schema merge
- plugin-repos/starlark-apps/pixlet_renderer.py — config value sanitization
- web_interface/blueprints/api_v3.py — timing key separation, safe manifest updates

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): use manifest filename field for .star downloads

Tronbyte apps don't always name their .star file to match the directory.
For example, the "analogclock" app has "analog_clock.star" (with underscore).

The manifest.yaml contains a "filename" field with the correct name.

Changes:
- download_star_file() now accepts optional filename parameter
- Install endpoint passes metadata['filename'] to download_star_file()
- Falls back to {app_id}.star if filename not in manifest

Fixes: "Failed to download .star file for analogclock" error

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): reload tronbyte_repository module to pick up code changes

The web service caches imported modules in sys.modules. When deploying
code updates, the old cached version was still being used.

Now uses importlib.reload() when module is already loaded.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): use correct 'fileName' field from manifest (camelCase)

The Tronbyte manifest uses 'fileName' (camelCase), not 'filename' (lowercase).
This caused the download to fall back to {app_id}.star which doesn't exist
for apps like analogclock (which has analog_clock.star).

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* feat(starlark): extract schema during standalone install

The standalone install function (_install_star_file) wasn't extracting
schema from .star files, so apps installed via the web service had no
schema.json and the config panel couldn't render schema-driven forms.

Now uses PixletRenderer to extract schema during standalone install,
same as the plugin does.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* feat(starlark): implement source code parser for schema extraction

Pixlet CLI doesn't support schema extraction (--print-schema flag doesn't exist),
so apps were being installed without schemas even when they have them.

Implemented regex-based .star file parser that:
- Extracts get_schema() function from source code
- Parses schema.Schema(version, fields) structure
- Handles variable-referenced dropdown options (e.g., options = dialectOptions)
- Supports Location, Text, Toggle, Dropdown, Color, DateTime fields
- Gracefully handles unsupported fields (OAuth2, LocationBased, etc.)
- Returns formatted JSON matching web UI template expectations

Coverage: 90%+ of Tronbyte apps (static schemas + variable references)

Changes:
- Replace extract_schema() to parse .star files directly instead of using Pixlet CLI
- Add 6 helper methods for parsing schema structure
- Handle nested parentheses and brackets properly
- Resolve variable references for dropdown options

Tested with:
- analog_clock.star (Location field) ✓
- Multi-field test (Text + Dropdown + Toggle) ✓
- Variable-referenced options ✓

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): add List to typing imports for schema parser

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): load schema from schema.json in standalone mode

The standalone API endpoint was returning schema: null because it didn't
load the schema.json file. Now reads schema from disk when returning
app details via web service.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* feat(starlark): implement schema extraction, asset download, and config persistence

## Schema Extraction
- Replace broken `pixlet serve --print-schema` with regex-based source parser
- Extract schema by parsing `get_schema()` function from .star files
- Support all field types: Location, Text, Toggle, Dropdown, Color, DateTime
- Handle variable-referenced dropdown options (e.g., `options = teamOptions`)
- Gracefully handle complex/unsupported field types (OAuth2, PhotoSelect, etc.)
- Extract schema for 90%+ of Tronbyte apps

## Asset Download
- Add `download_app_assets()` to fetch images/, sources/, fonts/ directories
- Download assets in binary mode for proper image/font handling
- Validate all paths to prevent directory traversal attacks
- Copy asset directories during app installation
- Enable apps like AnalogClock that require image assets

## Config Persistence
- Create config.json file during installation with schema defaults
- Update both config.json and manifest when saving configuration
- Load config from config.json (not manifest) for consistency with plugin
- Separate timing keys (render_interval, display_duration) from app config
- Fix standalone web service mode to read/write config.json

## Pixlet Command Fix
- Fix Pixlet CLI invocation: config params are positional, not flags
- Change from `pixlet render file.star -c key=value` to `pixlet render file.star key=value -o output`
- Properly handle JSON config values (e.g., location objects)
- Enable config to be applied during rendering

## Security & Reliability
- Add threading.Lock for cache operations to prevent race conditions
- Reduce ThreadPoolExecutor workers from 20 to 5 for Raspberry Pi
- Add path traversal validation in download_star_file()
- Add YAML error logging in manifest fetching
- Add file size validation (5MB limit) for .star uploads
- Use sanitized app_id consistently in install endpoints
- Use atomic manifest updates to prevent race conditions
- Add missing Optional import for type hints

## Web UI
- Fix standalone mode schema loading in config partial
- Schema-driven config forms now render correctly for all apps
- Location fields show lat/lng/timezone inputs
- Dropdown, toggle, text, color, and datetime fields all supported

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): code review fixes - security, robustness, and schema parsing

## Security Fixes
- manager.py: Check _update_manifest_safe return values to prevent silent failures
- manager.py: Improve temp file cleanup in _save_manifest to prevent leaks
- manager.py: Fix uninstall order (manifest → memory → disk) for consistency
- api_v3.py: Add path traversal validation in uninstall endpoint
- api_v3.py: Implement atomic writes for manifest files with temp + rename
- pixlet_renderer.py: Relax config validation to only block dangerous shell metacharacters

## Frontend Robustness
- plugins_manager.js: Add safeLocalStorage wrapper for restricted contexts (private browsing)
- starlark_config.html: Scope querySelector to container to prevent modal conflicts

## Schema Parsing Improvements
- pixlet_renderer.py: Indentation-aware get_schema() extraction (handles nested functions)
- pixlet_renderer.py: Handle quoted defaults with commas (e.g., "New York, NY")
- tronbyte_repository.py: Validate file_name is string before path traversal checks

## Dependencies
- requirements.txt: Update Pillow (10.4.0), PyYAML (6.0.2), requests (2.32.0)

## Documentation
- docs/STARLARK_APPS_GUIDE.md: Comprehensive guide explaining:
  - How Starlark apps work
  - That apps come from Tronbyte (not LEDMatrix)
  - Installation, configuration, troubleshooting
  - Links to upstream projects

All changes improve security, reliability, and user experience.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): convert Path to str in spec_from_file_location calls

The module import helpers were passing Path objects directly to
spec_from_file_location(), which caused spec to be None. This broke
the Starlark app store browser.

- Convert module_path to string in both _get_tronbyte_repository_class
  and _get_pixlet_renderer_class
- Add None checks with clear error messages for debugging

Fixes: spec not found for the module 'tronbyte_repository'

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): restore Starlark Apps section in plugins.html

The Starlark Apps UI section was lost during merge conflict resolution
with main branch. Restored from commit 942663ab which had the complete
implementation with filtering, sorting, and pagination.

Fixes: Starlark section not visible on plugin manager page

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): restore Starlark JS functionality lost in merge

During the merge with main, all Starlark-specific JavaScript (104 lines)
was removed from plugins_manager.js, including:
- starlarkFilterState and filtering logic
- loadStarlarkApps() function
- Starlark app install/uninstall handlers
- Starlark section collapse/expand logic
- Pagination and sorting for Starlark apps

Restored from commit 942663ab and re-applied safeLocalStorage wrapper
from our code review fixes.

Fixes: Starlark Apps section non-functional in web UI

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): security and race condition improvements

Security fixes:
- Add path traversal validation for output_path in download_star_file
- Remove XSS-vulnerable inline onclick handlers, use delegated events
- Add type hints to helper functions for better type safety

Race condition fixes:
- Lock manifest file BEFORE creating temp file in _save_manifest
- Hold exclusive lock for entire read-modify-write cycle in _update_manifest_safe
- Prevent concurrent writers from racing on manifest updates

Other improvements:
- Fix pages_v3.py standalone mode to load config.json from disk
- Improve error handling with proper logging in cleanup blocks
- Add explicit type annotations to Starlark helper functions

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): critical bug fixes and code quality improvements

Critical fixes:
- Fix stack overflow in safeLocalStorage (was recursively calling itself)
- Fix duplicate event listeners on Starlark grid (added sentinel check)
- Fix JSON validation to fail fast on malformed data instead of silently passing

Error handling improvements:
- Narrow exception catches to specific types (OSError, json.JSONDecodeError, ValueError)
- Use logger.exception() with exc_info=True for better stack traces
- Replace generic "except Exception" with specific exception types

Logging improvements:
- Add "[Starlark Pixlet]" context tags to pixlet_renderer logs
- Redact sensitive config values from debug logs (API keys, etc.)
- Add file_path context to schema parsing warnings

Documentation:
- Fix markdown lint issues (add language tags to code blocks)
- Fix time unit spacing: "(5min)" -> "(5 min)"

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): critical path traversal and exception handling fixes

Path traversal security fixes (CRITICAL):
- Add _validate_starlark_app_path() helper to check for path traversal attacks
- Validate app_id in get_starlark_app(), uninstall_starlark_app(),
  get_starlark_app_config(), and update_starlark_app_config()
- Check for '..' and path separators before any filesystem access
- Verify resolved paths are within _STARLARK_APPS_DIR using Path.relative_to()
- Prevents unauthorized file access via crafted app_id like '../../../etc/passwd'

Exception handling improvements (tronbyte_repository.py):
- Replace broad "except Exception" with specific types
- _make_request: catch requests.Timeout, requests.RequestException, json.JSONDecodeError
- _fetch_raw_file: catch requests.Timeout, requests.RequestException separately
- download_app_assets: narrow to OSError, ValueError
- Add "[Tronbyte Repo]" context prefix to all log messages
- Use exc_info=True for better stack traces

API improvements:
- Narrow exception catches to OSError, json.JSONDecodeError in config loading
- Remove duplicate path traversal checks (now centralized in helper)

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

* fix(starlark): logging improvements and code quality fixes

Logging improvements (pages_v3.py):
- Add logging import and create module logger
- Replace print() calls with logger.warning() with "[Pages V3]" prefix
- Use logger.exception() for outer try/catch with exc_info=True
- Narrow exception handling to OSError, json.JSONDecodeError for file operations

API improvements (api_v3.py):
- Remove unnecessary f-strings (Ruff F541) from ImportError messages
- Narrow upload exception handling to ValueError, OSError, IOError
- Use logger.exception() with context for better debugging
- Remove early return in get_starlark_status() to allow standalone mode fallback
- Sanitize error messages returned to client (don't expose internal details)

Benefits:
- Better log context with consistent prefixes
- More specific exception handling prevents masking unexpected errors
- Standalone/web-service-only mode now works for status endpoint
- Stack traces preserved for debugging without exposing to clients

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 19:44:12 -05:00