feat(web): weekly automatic updates with health check and rollback (#581)

* feat(web): weekly automatic updates with health check and rollback

A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.

- Pre-update checks skip (and report) instead of forcing: local edits or
  commits, merge/live rebase, no upstream, low disk, missing health check, or
  a version that was already rolled back. An abandoned rebase (HEAD back on a
  branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
  reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
  file, restarts the services from its own cgroup, requires them to come up
  and stay up, and otherwise resets to the previous commit and reinstalls the
  previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
  (as root) installs the two units from the repo templates for the web user.
  first_time_install.sh installs them too and takes --enable-auto-update /
  LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
  rollbacks raise an Overview banner and show under the toggle.

Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(auto-update): address static-analysis findings

- Replace the subprocess.CompletedProcess the verifier fabricated for a
  command that could not start with a plain namedtuple; nothing is executed
  there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
  calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): CI failures on Linux

- Keep the setup result when chown fails. CI runs as a non-root user, where
  chown to the web user raises; that discarded the result file, so the
  General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
  snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): address review feedback

- Health check: a failed restart command no longer lets the check run
  against the still-running old process; it counts as a failure (and after a
  rollback, as a failed rollback). An unreadable restart count is never
  treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
  keeping mode and owner, so a running config watcher never reads a
  truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
  setup refuses folder names systemd would reinterpret (%, quotes,
  backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): keep error detail in the status route's 500

test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): dismiss route rejects non-object JSON with 400

A JSON array or scalar body made `.get('alert_id')` raise, returning 500.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): let the app-wide handler answer status-route errors

CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-15 10:58:57 -04:00
committed by GitHub
co-authored by Claude Opus 5
parent d01da3bd9f
commit 869e36fb2f
25 changed files with 2938 additions and 192 deletions
+277
View File
@@ -0,0 +1,277 @@
#!/usr/bin/env python3
"""Check that an automatic LEDMatrix update left the device working; roll it back if not.
Started by the web interface's weekly updater (web_interface/auto_update.py)
through ledmatrix-update-verify.service, right after it pulls new code. It has
to run outside the web service: checking the update means restarting that
service, and a check running inside it would be killed by its own restart.
It must not be the code it is checking, either. The updater copies this file
to data/auto_update_verifier.py *before* pulling and the unit runs that copy,
so a broken update cannot break its own rollback. Standard library only for
the same reason: the rollback cannot depend on packages the update changed.
The updater leaves data/auto_update_pending.json:
{"status": "pending", "old_head": ..., "new_head": ...,
"display_was_active": bool, "dependency_failures": [...], "created_at": ...}
This moves its status to "verifying" and then to one of "success",
"rolled_back" or "rollback_failed", with "reason" and "detail" saying why.
The web interface reports that outcome and raises a banner for anything but
success.
"""
import json
import os
import subprocess # nosec B404 - list-form argv only, no shell # nosemgrep
import sys
from collections import namedtuple
import tempfile
import time
import traceback
import urllib.error
import urllib.request
from pathlib import Path
PENDING_NAME = 'auto_update_pending.json'
REQUIREMENT_FILES = ('requirements.txt', 'web_interface/requirements.txt')
WEB_HEALTH_URL = 'http://127.0.0.1:5000/api/v3/system/version'
#: How long the services get to come up after a restart...
HEALTH_TIMEOUT_SECONDS = 180
#: ...and how long they must then stay up. Restart=on-failure makes a crash
#: loop look healthy between attempts, so a single "is-active" proves nothing.
STABLE_SECONDS = 45
POLL_SECONDS = 5
PIP_TIMEOUT_SECONDS = 600
#: sudoers matches the exact command line, so bash is named by path, the same
#: candidates src/common/permission_utils.install_requirements_file tries.
BASH_CANDIDATES = ('/usr/bin/bash', '/bin/bash')
#: What a command that could not run at all reports: its callers only read
#: these three fields, the same ones a completed subprocess has.
_Failed = namedtuple('_Failed', 'returncode stdout stderr')
def pending_path(project_root):
return Path(project_root) / 'data' / PENDING_NAME
def read_pending(path):
try:
with open(path, 'r', encoding='utf-8') as f:
data = json.load(f)
return data if isinstance(data, dict) else None
except (OSError, ValueError):
return None
def write_pending(path, data):
path = Path(path)
path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp = tempfile.mkstemp(dir=str(path.parent), prefix='.auto_update_pending_')
try:
with os.fdopen(fd, 'w', encoding='utf-8') as f:
json.dump(data, f, indent=2)
os.replace(tmp, path)
except BaseException:
try:
os.unlink(tmp)
except OSError:
pass
raise
def _web_responds(url=WEB_HEALTH_URL):
try:
with urllib.request.urlopen(url, timeout=5) as resp: # nosec B310 - fixed loopback URL
return resp.status == 200
except (urllib.error.URLError, OSError, ValueError):
return False
def _short(sha):
return (sha or 'unknown')[:7]
class Verifier:
def __init__(self, project_root, run=subprocess.run, sleep=time.sleep,
clock=time.monotonic, web_responds=_web_responds, log=None):
self.project_root = Path(project_root)
self.pending_file = pending_path(project_root)
self.run = run
self.sleep = sleep
self.clock = clock
self.web_responds = web_responds
self.log = log or (lambda msg: print(f'[auto-update-verify] {msg}', flush=True))
def _run(self, args, timeout=60):
try:
return self.run(args, cwd=str(self.project_root), capture_output=True,
text=True, timeout=timeout)
except (subprocess.SubprocessError, OSError) as e:
return _Failed(returncode=1, stdout='', stderr=str(e))
# -- services ---------------------------------------------------------
def service_active(self, unit):
return self._run(['systemctl', 'is-active', unit], timeout=10).stdout.strip() == 'active'
def restart_count(self, unit):
out = self._run(['systemctl', 'show', '-p', 'NRestarts', '--value', unit],
timeout=10).stdout.strip()
return int(out) if out.isdigit() else None
def restart(self, unit):
result = self._run(['sudo', '-n', 'systemctl', 'restart', f'{unit}.service'], timeout=90)
if result.returncode != 0:
self.log(f'restarting {unit} failed: {(result.stderr or "").strip()}')
return result.returncode == 0
def restart_services(self, display):
"""Restart what should be running. False if any restart command failed."""
ok = True
# A display the user had stopped stays stopped.
if display:
ok = self.restart('ledmatrix') and ok
return self.restart('ledmatrix-web') and ok
def wait_healthy(self, display):
"""None once the services are up and stay up, else what went wrong."""
deadline = self.clock() + HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS
healthy_since = baseline = None
web = disp = False
count_known = True
while self.clock() < deadline:
web = self.web_responds()
disp = self.service_active('ledmatrix') if display else True
restarts = self.restart_count('ledmatrix') if display else None
# Without a restart count a crash loop looks healthy between
# attempts, so an unreadable count never counts as stable.
count_known = not display or restarts is not None
if web and disp and count_known and (healthy_since is None or restarts == baseline):
if healthy_since is None:
healthy_since, baseline = self.clock(), restarts
elif self.clock() - healthy_since >= STABLE_SECONDS:
return None
else:
healthy_since = None
self.sleep(POLL_SECONDS)
problems = []
if not web:
problems.append('the web interface did not respond')
if not disp:
problems.append('the display service did not stay running')
if web and disp and not count_known:
problems.append("the display service's restart count could not be read")
return '; '.join(problems) or 'the display service kept restarting'
# -- rollback ---------------------------------------------------------
def changed_requirements(self, old, new):
result = self._run(['git', 'diff', '--name-only', old, new])
# If the diff is unavailable, reinstall both rather than guess.
changed = set(result.stdout.split()) if result.returncode == 0 else set(REQUIREMENT_FILES)
return [rel for rel in REQUIREMENT_FILES if rel in changed]
def install_requirements(self, rel):
wrapper = self.project_root / 'scripts' / 'fix_perms' / 'safe_pip_install.sh'
req = self.project_root / rel
if not req.exists():
return True
for bash in BASH_CANDIDATES:
result = self._run(['sudo', '-n', bash, str(wrapper), str(req)],
timeout=PIP_TIMEOUT_SECONDS)
if result.returncode == 0:
return True
return False
def rollback(self, pending):
"""Reset to the previous commit and its dependencies. Returns (ok, detail)."""
old, new = pending.get('old_head'), pending.get('new_head')
if not old:
return False, 'the commit to roll back to is unknown'
requirements = self.changed_requirements(old, new) if new else list(REQUIREMENT_FILES)
# Safe to --hard: the updater refuses to run with local edits to
# tracked files, so the only thing this discards is the update.
result = self._run(['git', 'reset', '--hard', old], timeout=120)
if result.returncode != 0:
return False, (f'"git reset --hard {old}" failed: '
f'{(result.stderr or result.stdout or "").strip()}')
failed = [rel for rel in requirements if not self.install_requirements(rel)]
if failed:
return True, ('reinstalling the previous dependencies from ' + ', '.join(failed)
+ ' failed; run Install Base Requirements from the Tools tab')
return True, ''
# -- the check itself -------------------------------------------------
def _finish(self, pending, status, reason=None, detail=None):
pending.update({'status': status, 'reason': reason, 'detail': detail or None,
'finished_at': time.time()})
write_pending(self.pending_file, pending)
self.log(' '.join(p for p in (status, reason or '', detail or '') if p))
def verify(self):
pending = read_pending(self.pending_file)
if not pending or pending.get('status') != 'pending':
self.log('no update is waiting to be verified')
return 0
pending['status'] = 'verifying'
write_pending(self.pending_file, pending)
display = bool(pending.get('display_was_active'))
dependency_failures = pending.get('dependency_failures') or []
if dependency_failures:
# Never restart onto code whose packages did not install.
reason = 'installing its dependencies failed (' + ', '.join(dependency_failures) + ')'
elif not self.restart_services(display):
# The old process may still be answering; checking it would pass
# an update that never started.
reason = 'restarting the services failed'
else:
reason = self.wait_healthy(display)
if reason is None:
self._finish(pending, 'success')
return 0
self.log(f'update to {_short(pending.get("new_head"))} is unhealthy ({reason}); '
f'rolling back to {_short(pending.get("old_head"))}')
ok, detail = self.rollback(pending)
if not ok:
self._finish(pending, 'rollback_failed', reason, detail)
return 1
still = (self.wait_healthy(display) if self.restart_services(display)
else 'restarting the services failed')
if still:
self._finish(pending, 'rollback_failed', reason,
f'still unhealthy after rolling back: {still}'
+ (f'; {detail}' if detail else ''))
return 1
self._finish(pending, 'rolled_back', reason, detail)
return 0
def main(argv):
if len(argv) != 2:
print('usage: auto_update_verify.py PROJECT_ROOT', file=sys.stderr)
return 2
verifier = Verifier(Path(argv[1]))
try:
return verifier.verify()
except Exception as e:
traceback.print_exc()
# Whatever happened, the web interface must not be left thinking the
# check is still running.
try:
pending = read_pending(verifier.pending_file) or {}
if pending.get('status') in ('pending', 'verifying'):
pending.update({'status': 'rollback_failed', 'reason': 'the health check crashed',
'detail': str(e), 'finished_at': time.time()})
write_pending(verifier.pending_file, pending)
except OSError:
pass
return 1
if __name__ == '__main__':
sys.exit(main(sys.argv))