Commit Graph
9 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5 d1e821c625 fix(web): harden, polish and optimize the web UI per the Sept 2026 audit (#568)
* fix(web): harden, polish and optimize the web UI per the September 2026 audit

Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20).

Implementation integrity (P0)
- app.css now defines every utility class the templates and JS use,
  including .hidden, so the ~145 JS show/hide toggles work. Button reset,
  and base component rules (.btn, .form-control) wrapped in :where() so
  utility classes on the same element win. New static-audit test fails
  when a used utility class has no rule.

Accessibility
- Focus rings render (the old ring rule referenced undefined variables);
  one :focus-visible outline everywhere; skip link; labelled nav landmarks.
- Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap,
  Escape, focus return, applied to every modal.
- Named icon-only buttons and labelled ~70 form fields.
- Toasts announced once; errors persist >= 10s; one showNotification.
- Captive WiFi page: live region, timeouts, dark mode, 16px inputs.

Performance (Pi Zero 2 W)
- SSE streams and tab timers pause when hidden or off-tab; the display
  stream only runs while a preview is visible. app-shell.js deferred.
- Widget scripts served as one versioned bundle (/assets/widgets.js):
  52 -> 21 script tags, 66 -> 35 requests on first load.
- Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS
  1358 KB -> 291 KB on the wire. SSE untouched.

Theming and responsive
- File managers, form fields and Fonts upload on theme tokens; bare
  inputs themed in dark mode; no more white surfaces.
- No horizontal overflow at 375px on any tab; 44px touch targets on
  coarse pointers; reduced-motion respected; header title truncates.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings on #568

- json-file-manager: focus-trap releases kept in a Map (no dynamic
  property access or delete; no value-returning forEach callback)
- notification / schedule-picker: style and day-label lookups via Map
- app.js: move the pending-queue assignment out of the expression
- diff_viewer / error_handler: named function declarations instead of
  arrow consts

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: check the OAuth widget ships in the widget bundle

base.html no longer tags widget scripts one by one; they load through
/assets/widgets.js. Assert the page requests the bundle and the bundle
contains google-oauth.js, which is what the test was protecting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): address review feedback on #568

- widget bundle version fingerprints every file (name, mtime_ns, size)
- gzip fallback appends Accept-Encoding to an existing Vary header
- dialog helper: releasing a non-top dialog no longer moves focus out of
  the dialog the user is in
- labels: file-upload targets its file input; fallback config fields get
  label for/id pairs; native color input has a fallback name
- utility audit also reads class names inside bound :class expressions

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give the native color-picker input an accessible name

CodeRabbit flagged this on PR #568 as an outside-diff finding (never
posted inline, so it was missed in the round of fixes that addressed
the other 6 review comments). The <input type="color"> only carried a
title attribute; screen readers don't reliably announce title, and
there's no other label naming the control when showHexInput is false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings in app-shell.js

- drop the unused catch binding on the SSE JSON parse
- move the pending-notification queue assignment out of the expression

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): contain plugin widgets/ dir and bound style-editor retries

From CodeRabbit review on #568 (code that arrived with the main merge):
- serve_plugin_widget resolves widgets/ with resolve_under before
  resolving the manifest script under it, so a symlinked widgets
  directory can't become the containment base (CWE-22). New test.
- style-editor init stops polling after ~10s when the widget never
  registers and leaves the plain fallback fields in place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 09:42:24 -04:00
ChuckandClaude Opus 5 5137e86d16 feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab

PR #544's change, ported onto the api_v3 package split (#553). Identical
behaviour; only the placement of the new code differs.

The original added 508 lines to web_interface/blueprints/api_v3.py, which #553
deletes, so every hunk of it would conflict irreconcilably. Ported by AST:
26 new top-level items sorted to where the split puts each kind --

  __init__.py   2 imports, 11 constants, 7 helpers
  starlark.py   4 routes  (/starlark/editor/{apps,status,start,stop})
  misc.py       2 routes  (/integrations/mqtt-bridge{,/config})

Everything outside api_v3.py -- the Tools partial, the installer scripts, the
JS tests -- applied unchanged.

Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is
regenerated to match, which is exactly what test_api_v3_url_map.py is designed
to make you do when routes are added.

Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR
ships could not be run here -- node is not installed on this machine -- so
test/js/dom/test_tools_sections.js is unverified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST

The AST-based port of #544 onto the api_v3 package split dropped `time`
from starlark.py's import list. start_pixlet_editor() and
stop_pixlet_editor() both call time.time()/time.sleep() directly, so
every start (NameError building `state['started_at']`) and every stop
that has to wait out the EXIT trap crashed with a 500. No test caught
it because the route's own tests mock subprocess.Popen but never
actually invoked it before now.

Also carries over #544's later fix that this port branched before:
env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an
operator who had already pinned PIXLET_EDITOR_HOST to loopback,
forcing the unauthenticated `pixlet serve` process onto the LAN
regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as
api_v3.starlark.py's siblings already do for _pkg-owned names.

Both fixes route the shared _pkg.time reference the rest of the
package's route modules already use for anything a test might need to
patch, rather than a bare `import time` local to this file.

Ported the existing regression test from #544
(TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's
module layout (web_interface.blueprints.api_v3.starlark instead of the
old monolithic api_v3 module), which is what caught the NameError.

Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this
branch and on origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to
this change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): clear the six lint errors this rebase introduced

All six were introduced by rebasing this branch onto the merged blueprint
split, not by the split itself. Confirmed by diffing pyflakes output against
main with line numbers normalised -- everything else it reports is present on
main too and is the package's deliberate re-export pattern.

starlark.py used _STARLARK_APPS_DIR three times without importing it (F821).
The rebase resolved an import-list conflict as a union of both sides, and that
symbol was on neither side of the conflict hunk, so it was silently lost. It is
defined in __init__.py and is now imported like its neighbours. This was the
only one of the six that would fail at runtime rather than merely lint.

__init__.py imported contextlib twice (F811): the cherry-pick added one next to
the existing import. Removed the duplicate; the original at line 19 is used.

__init__.py imported signal purely to re-export it to starlark.py, so pyflakes
saw it as unused (F401). signal is stdlib and does not need routing through the
blueprint package, so starlark.py imports it directly and __init__.py no longer
does. contextlib stays re-exported because this module genuinely uses it.

_read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule
this module imports at the bottom for its route side effects (F811). Renamed to
`settings`, with a comment saying why, since the name is otherwise the obvious
one to reach for.

Verified: pyflakes now reports nothing on this branch that main does not, the
package imports, all nine route modules load, and 117 routes register, matching
the pinned URL-map snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply

Two CodeRabbit findings on the bridge settings endpoint, both of which returned
200 while doing something other than what the caller asked.

`request.get_json(silent=True) or {}` turned a missing or unparseable body --
and the JSON literals null, [] and false -- into an empty dict, which then
satisfied the isinstance(data, dict) guard on the very next line. The guard was
there to reject exactly those bodies. Dropping the `or {}` lets None fail it.

The same `or {}` on /errors/clear is left alone: its docstring documents the
body as optional, so an absent body legitimately means "use the defaults". The
difference is that saving settings has nothing sensible to do with no body.

`if data.get('clear_password'):` accepted any truthy value, and the string
"false" is truthy in Python -- so a client echoing the field back as a string
wiped a password it meant to keep. Now coerced through the package's existing
_coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else.

test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus
a missing body, and clear_password across truthy and falsy spellings. Verified
against the unfixed code -- reverting the body guard fails 5, reverting the
coercion fails 3.

Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials
when TLS is off (CWE-319). That is a policy decision about the feature rather
than a defect -- unencrypted MQTT on a trusted LAN is common and often
deliberate -- so it is raised on the PR for a maintainer call instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: work through the remaining review findings on the editor and bridge

allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses
the network in cleartext. Refused now rather than merely warned about -- but
refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal
deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults
false, and is coerced like the other booleans so the string "false" cannot
switch the guard off.

starlark.py:796 -- the supported service runs Flask threaded, so two start
requests could each see running=False, each launch an editor, and the second
state write replace the first PID, orphaning a process that holds the display
down with nothing recording it. The check-launch-write sequence now takes a
module-level lock.

starlark.py:848 -- if the state write failed the route returned success with an
editor running and no PID recorded: status and stop both reported no session
while the display stayed down until the timeout expired. It now terminates the
process group and returns an error.

starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so
nothing hands the display back, yet the response said "the display is
restarting". After an escalation the display is now restarted explicitly, and a
failure to do so returns an error naming the manual step instead of a success.

pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no
timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the
failure lands before the display is stopped rather than after.

pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface
address are three cases, not two. Collapsing the last two printed a URL saying
"localhost" whenever PIXLET_EDITOR_HOST named a LAN address.

tools.html:1254 -- escHtml does not encode single quotes, and the app id was
interpolated into an inline onclick="startPixletEditor('...')", so a directory
containing an apostrophe could break out of the JS string and run script. The
handler binds with addEventListener and reads the id from dataset, where it is
only ever parsed as an HTML attribute.

Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in
in both directions. The tools DOM suite gains three guards asserting the edit
buttons carry no inline onclick and pass the id via dataset -- those need jsdom
and did not run here, so CI verifies them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): log the traceback on the editor state-write failure

The 848 fix answers 500 when the session state cannot be written, and logged
that at error level -- but without exc_info, so the traceback never reached the
log. test_web_error_detail.py guards exactly this: a handler returning 5xx must
write an error-level record *with* the traceback and return the sanitized
detail, because checking that merely something was logged is too weak.

Caught by Core unit tests on the previous commit, not locally: the guard parses
every module under web_interface/blueprints/api_v3 as one source, so it only
fires once the whole package is read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 11:32:12 -04:00
ChuckandClaude Opus 5 f90638a9ec feat(web): show available memory in Tools diagnostics (#500)
* feat(web): show available memory in Tools diagnostics

System Diagnostics reported memory as used-percent plus used/total GB.
Neither distinguishes a healthy board from one about to fail, because
page cache counts as used and is reclaimable on demand -- a Pi can read
70% used and be fine, or read the same and be minutes from trouble.

MemAvailable is the kernel's own estimate of what a new allocation can
actually obtain, and it is the number that tracked the failure on a 1GB
Pi 3B+: healthy running sat above 500MB, and the crash came at 73MB. By
that point fork() was failing, so sshd could not spawn a session and
systemd could not respawn the display, while the kernel carried on
answering pings at 0% loss. Used-percent gave no warning at any point on
the way there; available memory fell steadily for hours.

/api/v3/system/status now returns memory_available_mb from
psutil.virtual_memory().available, and Tools renders it as its own tile,
coloured against the thresholds that failure implies: red under 150MB,
amber under 300MB, green above.

The existing memory tile is left alone -- used/total is still what you
want when sizing a workload; this answers the different question of how
much room is left right now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): round available memory once, so the tile agrees with itself

The colour was classified from the raw value while the label was rounded,
and the API sends one decimal place. At the boundaries the two disagreed:
149.6 rendered as "150 MB" in red, and 299.6 as "300 MB" in amber -- each
contradicting the threshold its own colour claims to apply ("red under
150MB"). A reader checking the tile against the documented thresholds would
conclude the readout was broken.

Rounding once and using that number for both restores agreement. It moves
those two boundary cases up a band, which does not matter: the thresholds
come from a measured failure at 73MB, so which side of the line a spare
0.4MB falls on carries no information. The tile agreeing with itself does.

Null handling and the thresholds themselves are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 08:40:28 -04:00
Chuck fc25a70d75 fix(web): make the update button work on branches without tracking (#443)
Reported from a pi whose checkout sat on a local branch:

    git pull failed (returncode=1): There is no tracking information for
    the current branch. Please specify which branch you want to rebase
    against.

The Tools tab reported that as "Update failed; check logs for details",
which tells the user nothing they can act on, and the underlying git
message never reached the UI at all.

A branch with no upstream is easy to end up on — checking one out by
name, restoring a backup, or following a guide that names a branch — and
until now it left the update button permanently broken with no way out
except SSH.

resolve_pull_command() now decides how to pull:
  - upstream set                  -> git pull --rebase, as before
  - no upstream, origin/<branch>  -> git pull --rebase origin <branch>,
    then attach tracking so the next update is a plain pull
  - no upstream, no remote branch -> an error naming the branch and
    pointing at Switch branch
  - detached HEAD                 -> says so, rather than failing obscurely

That resolution happens BEFORE the stash. Previously the handler stashed
local changes and then discovered it could not pull, putting the user's
work away for an update that was never going to run.

Failures now surface git's own message instead of "check logs".

Adds a branch picker to the Tools tab, backed by GET
/system/git-branches (local + remote-only) and a checkout_branch action.
Switching attaches tracking, so Pull Latest works afterwards. Branch
names are validated against a strict pattern before reaching a subprocess
argument list.

Local edits block a checkout, as they should. Rather than a truncated
one-line error, the response carries git's full list of blocking files
and a can_retry_with_stash flag; the UI then offers "Stash and switch" as
an explicit choice. Stashing is never done unasked — putting someone's
edits away without consent is worse than refusing the switch.

Verified on the pi that produced the report: on its untracked 'audit'
branch the update now returns the actionable message, git-info reports
upstream='' and can_pull=false, and an injected branch name is rejected.
27 tests build real git repositories and cover each path, including the
stash route that could not be exercised safely on the device.
2026-08-07 13:27:43 -04:00
ChuckandClaude Opus 4.8 bd9f461f70 Add system diagnostics, power controls, and WiFi radio toggle to Tools tab (#389)
* Add system diagnostics, power controls, and WiFi radio toggle to Tools tab

Expands the web UI Tools tab with safe, purpose-built controls so users can
manage the Pi without SSHing in, instead of an arbitrary-command terminal
(the web UI has no auth and CSRF is disabled, so a shell would be unsafe).

- System Diagnostics card: renders the existing but previously-unused
  GET /api/v3/system/status endpoint (CPU, memory, temp, disk, uptime),
  with a manual refresh and a 10s poll.
- System Power section: reboot/shutdown buttons wired to the existing
  reboot_system / shutdown_system actions, behind a confirm step, with a
  dedicated powerAction() helper that treats the dropped connection as the
  expected "going offline" outcome rather than an error.
- Network Radio section: WiFi on/off toggle backed by new
  GET/POST /api/v3/wifi/radio endpoints and WiFiManager.set_wifi_radio() /
  get_wifi_radio_state(). Disabling WiFi is refused unless a wired
  connection is present (reusing the existing lockout guards), with an
  explicit force-off confirmation for advanced users.

No new privileged commands: uses nmcli radio wifi on|off (already
sudo-allowlisted) and the existing reboot/poweroff grants, so the
sudoers-alignment guard test stays green. Bluetooth toggle intentionally
omitted since the installer removes the BlueZ stack for LED timing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019X3he87Ggr1qt8y7mnFeuE

* Harden WiFi radio endpoint and diagnostics poll (review feedback)

Addresses code review feedback on the Tools tab additions:

- api_v3.py: parse the `enabled` POST field with the same string-aware
  coercion as `force`. A plain bool() cast turned {"enabled":"false"} into
  True (enabling instead of disabling) for any non-UI API caller.
- wifi_manager.set_wifi_radio() now returns a reason code alongside
  (success, message); the /wifi/radio error response includes it. The Tools
  UI only shows the force-off confirmation when reason == 'no_ethernet', so a
  genuine nmcli failure surfaces its real error instead of a misleading
  "no wired connection" prompt that would just retry into the same failure.
- tools.html: gate the 10s diagnostics poll on panel visibility
  (document.hidden / offsetParent), so switching to another tab stops the
  recurring /api/v3/system/status calls instead of churning the Pi off-screen.
  The initial load and manual Refresh remain unconditional.

Verified: {"enabled":"false"} now disables (refused w/ reason:no_ethernet),
{"enabled":"true"} enables; inline JS passes node --check; py_compile clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019X3he87Ggr1qt8y7mnFeuE

* Narrow WiFi-disable fallback except and add stack trace (review)

Addresses a CodeRabbit nitpick: the fallback handler in set_wifi_radio()'s
disable path caught bare Exception and logged without a traceback. Narrow it to
(OSError, subprocess.SubprocessError) — the errors subprocess.run realistically
raises — and log with exc_info=True for full context on the Pi. Anything
genuinely unexpected now propagates to the endpoint's outer handler (500),
matching the codebase's specific-exception convention.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019X3he87Ggr1qt8y7mnFeuE

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-09 09:22:11 -04:00
ChuckandClaude Opus 4.8 3b93024993 feat: activate dormant plugin health/metrics subsystem and surface it in the web UI (#388)
* feat(plugin-system): activate dormant plugin health & metrics subsystem

PluginManager shipped a fully-built health tracker, resource monitor and
circuit breaker that were never instantiated (health_tracker/resource_monitor
were left as None), so the circuit breaker never engaged and the existing
health/metrics API routes always returned "not available".

- DisplayController now wires a PluginHealthTracker and PluginResourceMonitor
  onto the plugin manager, enabling the circuit breaker (a repeatedly-failing
  plugin's update() is skipped after consecutive failures, then retried after
  a cooldown) and per-plugin execution-time metrics. Both persist to the
  shared cache.
- load_plugin() now validates each plugin's config against its JSON schema in
  a strictly warn/degrade-only way: a violation logs a warning and flags the
  plugin degraded in the health tracker, but never changes whether the plugin
  loads or its pass/fail behaviour. Adds PluginHealthTracker.set_degraded(),
  which never touches the circuit breaker.
- ResourceMonitor CPU/memory sampling now reuses a cached psutil.Process and
  reads cpu_percent(interval=None), so monitoring no longer blocks ~100ms per
  call on the display loop's update path.
- Fix DiskCache.get() raising TypeError for max_age=None ("never expires"),
  which silently discarded persisted plugin health/metrics on read and thus
  broke cross-process and post-restart surfacing.
- Fix two dead PluginManager helpers that called non-existent tracker methods.

Tests: new test_resource_monitor, test_plugin_health,
test_plugin_manager_schema_soft; extended test_cache_manager and
test_display_controller.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

* feat(web-ui): surface plugin health, metrics and load state

With the health/metrics subsystem now active in the display service, expose it
in the web UI (which runs as a separate process from the display loop):

- Wire a health tracker / resource monitor backed by the shared on-disk cache
  into the web process so /api/v3/plugins/health and /plugins/metrics read the
  data the display service persists.
- Build those route responses per installed plugin id (the tracker's in-memory
  view is empty in a fresh web process) so cross-process data is included.
- Add state + error_info to /plugins/installed entries so the UI can show why a
  plugin isn't running instead of just loaded:false.
- Add a "Plugin Health" panel to the Tools page (circuit status, avg/max update
  time, update count, last error) plus PluginAPI.getPluginMetrics().

Tests: route-level tests for the health/metrics endpoints in test_web_api.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

* fix(plugin-metrics): refresh cross-process health/metrics reads; type hints

Addresses CodeRabbit review on #388:

- Major: the web process's health/resource trackers cached the first persisted
  read in an in-memory dict (and the CacheManager memory tier held max_age=None
  entries indefinitely), so a long-lived web process showed the first snapshot
  and never reflected the display service's later updates. Add an opt-in
  force_reload path (get_health_summary/get_health_state/_load_health_state and
  get_metrics_summary/get_metrics) that bypasses the in-memory copy and, via a
  new memory_ttl passthrough on CacheManager.get, the cache manager's memory
  tier — so each /plugins/health and /plugins/metrics poll reads fresh persisted
  state. Default behaviour (force_reload=False) is unchanged for the display
  process and existing callers.
- Minor: DiskCache.get type hint is now Optional[int] with the None ("never
  expires") semantics documented, matching MemoryCache.get.

Tests: new force_reload staleness cases in test_plugin_health and
test_resource_monitor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-09 07:54:18 -04:00
ChuckandClaude Sonnet 5 9b2f02681d feat(web-ui): detect and surface Raspberry Pi under-voltage/throttling (#383)
* feat(web-ui): detect and surface Raspberry Pi under-voltage/throttling

Adds a vcgencmd get_throttled check to the system-status SSE stream and
surfaces it in the web UI:

- A header badge (next to CPU/Memory/Temp) that stays hidden when healthy,
  turns red when under-voltage/throttling is happening right now, and
  yellow if it happened earlier this session but has since cleared.
- A dismissible top banner (same pattern as the update-available banner)
  that appears while under-voltage/throttling is actively occurring, with
  guidance to check the power supply. Re-appears on a fresh occurrence
  even if a previous one was dismissed.
- A "Power Supply" card on the Overview tab alongside CPU/Memory/Temp/
  Display Status.

Motivated by a real device showing intermittent brightness flicker that
turned out to be ~1 under-voltage event every 30-90s (visible live via
`vcgencmd get_throttled` and dmesg's "Undervoltage detected!" messages) --
there was no way to see this from the web UI, only by SSHing in.

Returns None on non-Pi platforms (no vcgencmd on PATH), matching the
existing guard pattern used for the CPU temperature read.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web-ui): add power supply diagnostics detail to Tools tab

The header badge/banner/Overview card added in the previous commit only
show a collapsed "is it bad right now" signal. This adds a "Power Supply"
section to the Tools tab with the full 8-flag breakdown (under-voltage,
throttled, freq-capped, soft-temp-limit -- each split into "right now" vs
"occurred since boot") for actually troubleshooting a recurring issue,
plus a pointer to the README's power supply sizing guidance when something
is or was flagged.

Reuses the existing stats SSE stream (window.statsSource) rather than
adding a new endpoint -- the same payload already drives the header/
banner/Overview card, so this just listens for it too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web-ui): address review findings on power supply monitoring

- Overview card left its "--" placeholder forever on non-Pi platforms since
  the update only handled a truthy data.power. Now explicitly renders "Not
  available" with a neutral icon/color for that case.
- vcgencmd failures logged at debug, invisible in default remote log
  output. Bumped to warning to match the nearby systemctl failure logging.
- The banner/badge/card and the Tools summary line only looked at
  under_voltage_now/throttled_now (+ occurred), silently ignoring
  freq_capped_now/occurred and soft_temp_limit_now/occurred from
  _get_power_status() -- a Pi that's actively soft-thermal-limited or
  frequency-capped showed a green "OK" everywhere except the detailed
  flag table buried in Tools. All four surfaces now fold all four "now"/
  "occurred" flags into the same active/occurred state.
- The banner text was hardcoded to "Under-voltage detected..." even when
  the actual active condition was throttling/freq-capping/thermal limiting.
  Added _activePowerConditionLabels() (shared, non-module global scope) to
  build the banner/tooltip text from whichever flags are actually set.

Skipped: TTL-caching _get_power_status() to avoid "multiplying forks
across browser tabs" -- that premise doesn't hold against this codebase.
_StreamBroadcaster (its own docstring says as much) already runs exactly
one shared generator per tick regardless of client count, identical to
the uncached cpu_temp file-read two lines above it; there's nothing to
multiply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web-ui): address second round of review findings on power monitoring

- updatePowerStatus's falsy-power branch only hid power-stat, leaving
  power-warning-banner visible with stale text if _get_power_status()
  fails transiently. Now hides the banner too and resets the dismissed
  flag, same as the "not active" branch.
- The Tools tab's status badge (and the pre-existing dirty/clean badge
  right next to it) build class names like bg-${color}-100/text-${color}-800
  at runtime. This project hand-rolls its own Tailwind-named utility
  classes in app.css rather than running a real Tailwind build, and the
  light-mode base rules for bg-red-100/bg-yellow-100/bg-green-100/
  text-red-800/text-yellow-800/text-green-800 were simply never defined --
  only some had dark-mode overrides, which are no-ops without a base rule
  in light mode. Added the missing light-mode bases plus the two missing
  dark-mode overrides (bg-yellow-100/bg-green-100), fixing both badges.
- The "occurred earlier" tooltip was a hardcoded string regardless of
  which flag(s) actually fired. Generalized _activePowerConditionLabels()
  to take a suffix ('_now' or '_occurred') and reused it for both tooltips.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* refactor(web-ui): move Power Supply status off the Overview tab

The Overview tab's "Power Supply" stat card duplicated what the Tools
tab's diagnostics section already shows (summary badge + full flag
breakdown), so drop the card and its now-dead JS rather than keep two
copies in sync. The header badge and warning banner (visible on every
page) are unaffected -- only the Overview-tab card is removed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-07 09:35:55 -04:00
639e1c3a93 fix(web): repair news ticker custom-feeds save for JSON path (#376)
* feat(install): surface root cause of web dependency install failures

install_dependencies_apt.py previously reported only which packages
failed, not why - the actual apt/pip error was discarded (apt) or
could scroll out of the on_error log tail (pip), leaving "Step 7:
Install web interface dependencies (line 915)" as the only visible
detail.

Capture command output for each install attempt and print a compact
DEPENDENCY INSTALLATION FAILURES summary with the last lines of error
output per package. Also run the installer with `python3 -u` for
real-time, correctly-ordered logging, and widen the on_error tail from
50 to 100 lines so the summary isn't cut off.

* fix(web): repair news ticker custom-feeds save for JSON path

The JS dotToNested() helper converts indexed form fields like
feeds.custom_feeds.0.name into a dict {'0': {name:...}} rather than a
proper array. The form-data path already had fix_array_structures() to
convert those dicts back to arrays before schema validation, but the
JSON path (used by all web-UI saves) never ran that fix, so saving any
custom feed produced a schema validation error: "Expected type array,
got object".

Add _fix_json_arrays() immediately after schema loading on the JSON
path, mirroring the existing fix_array_structures() logic.

Also fix custom-feeds.js getValue() to omit the logo key entirely when
no logo is present instead of returning logo:null, which would fail
schema validation (logo expects type object).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(install,tools): address PR 376 review findings

- first_time_install.sh: add _clone_rpi_rgb() wrapper so retry() cleans
  up any partial rpi-rgb-led-matrix-master dir before each clone attempt
- first_time_install.sh: use apt-get -o DPkg::Lock::Timeout=180 so apt
  handles lock contention natively instead of relying solely on flock TOCTOU check
- install_dependencies_apt.py: pass DPkg::Lock::Timeout=180 to apt-get
  install to avoid failing when unattended-upgrades holds the lock
- install_dependencies_apt.py: add type annotations to all public helpers
- api_v3.py: fix install_plugin_requirements to read plugin_manager from
  api_v3 blueprint attribute instead of the always-None module variable
- tools.html: loadGitInfo() now checks r.ok before parsing JSON and
  surfaces d.status === 'error' with the server's message in the panel

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools,api): address three additional review findings

- api_v3.py install_plugin_requirements: replace hardcoded plugin-repos
  fallback with config-driven resolution (plugin_system.plugins_directory),
  matching the pattern used elsewhere in the module
- api_v3.py _fix_json_arrays: recurse into converted and existing array
  elements when items.type is object, so nested numeric-keyed dicts inside
  array items are also normalized
- tools.html toolsAction: check r.ok before r.json() and recover
  gracefully from non-JSON error bodies (HTML 500 pages), consistent
  with the existing loadGitInfo guard

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 14:21:05 -04:00
6096a22c3d feat(web): add Tools tab and row address type display setting (#373)
* feat(web): add Tools tab and row address type setting

Adds a Tools/Utilities tab to the web interface with one-click
maintenance buttons that previously required SSH:
- Git status panel (branch, dirty state, recent commits)
- Pull latest (rebase) and force reset to origin/main
- Reinstall base requirements (pip, with output)
- Reinstall per-plugin requirements (pass/fail per plugin)
- Clear __pycache__ directories
- Quick-access restart for display and web services

Also exposes the hzeller row_address_type option (0–4) in the
Display settings tab. The backend already read this value from
config; the UI, API field list, and validation were missing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): address code review findings

- Add _GIT = shutil.which('git') alongside _SUDO/_JOURNALCTL; return
  503 in force_git_reset and get_git_info if git is unavailable
- Check git branch/status returncodes in get_git_info(); return a clear
  500 error instead of silently treating a failed run as a clean repo
- Cap pip stdout+stderr at 50 KB via _truncate_output() helper to
  avoid OOM on verbose dependency resolution or build failures
- Scrub embedded HTTPS credentials from remote_url via
  _scrub_git_remote_url() using urllib.parse before returning to UI
- Fix clear_pycache to track and report failed deletions separately
  instead of counting them as successes (removed ignore_errors=True,
  wrapped in try/except OSError)

Skipped: plugin_manager-vs-api_v3.plugin_manager (api_v3 is the
Blueprint object; accessing .plugin_manager on it would fail — module-
level variable is the correct pattern used throughout this blueprint);
pages_v3 broad-except (identical to every other _load_*_partial in the
file); base.html HTMX fallback (loadTabContent handles all tabs
generically; named fallbacks only exist for tabs needing JS re-init);
tools.html auth (pre-existing architectural decision — reboot/shutdown
on the same endpoint are also unauthenticated).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): resolve remaining PR review comments

- api_v3: use getattr(api_v3, 'plugin_manager', None) instead of the
  module-level plugin_manager (always None); app.py sets the blueprint
  attribute, not the module global, so the fallback to plugin-repos was
  always taken
- pages_v3: replace broad except Exception in _load_tools_partial with
  specific TemplateNotFound / OSError handlers and add [Pages V3][Tools]
  context prefix to log messages and error responses for easier Pi
  debugging
- base.html: add Tools tab branch to the HTMX-unavailable fallback block
  in loadTabContent so the tab loads gracefully via direct fetch if HTMX
  never initialises

Skipped: auth on execute_system_action — pre-existing app-wide design;
reboot/shutdown and all other system actions share the same exposure.
An app-level auth layer is the correct fix and is out of scope here.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): resolve second-pass review findings

- Wrap per-plugin subprocess.run in try/except TimeoutExpired/OSError so
  one plugin's failure appends a result entry and continues the loop
  rather than collapsing the whole batch into a 500
- Validate double_sided_copies divisibility against chain_length
  (horizontal axis) or parallel (vertical axis) after the range check;
  reads effective axis from the current request or stored config
- Exclude double_sided_fields from the generic key-merge loop so
  double_sided_enabled/copies/axis are never written as root-level keys
- Fix tools.html copy: "then restores the stash" removed — git_pull
  stashes changes but never pops them
- Check r.ok and d.status in loadGitInfo before building the panel;
  backend error messages now surface instead of silently showing a
  false-clean state

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): don't expose filesystem paths in OSError messages

CodeQL flagged str(exc) flowing into the JSON response for the
install_plugin_requirements action. Use exc.strerror instead, which
gives the OS error description ("No such file or directory",
"Permission denied") without the internal filesystem path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 12:19:54 -04:00