feat(web): weekly automatic updates with health check and rollback (#581)

* feat(web): weekly automatic updates with health check and rollback

A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.

- Pre-update checks skip (and report) instead of forcing: local edits or
  commits, merge/live rebase, no upstream, low disk, missing health check, or
  a version that was already rolled back. An abandoned rebase (HEAD back on a
  branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
  reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
  file, restarts the services from its own cgroup, requires them to come up
  and stay up, and otherwise resets to the previous commit and reinstalls the
  previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
  (as root) installs the two units from the repo templates for the web user.
  first_time_install.sh installs them too and takes --enable-auto-update /
  LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
  rollbacks raise an Overview banner and show under the toggle.

Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(auto-update): address static-analysis findings

- Replace the subprocess.CompletedProcess the verifier fabricated for a
  command that could not start with a plain namedtuple; nothing is executed
  there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
  calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): CI failures on Linux

- Keep the setup result when chown fails. CI runs as a non-root user, where
  chown to the web user raises; that discarded the result file, so the
  General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
  snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): address review feedback

- Health check: a failed restart command no longer lets the check run
  against the still-running old process; it counts as a failure (and after a
  rollback, as a failed rollback). An unreadable restart count is never
  treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
  keeping mode and owner, so a running config watcher never reads a
  truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
  setup refuses folder names systemd would reinterpret (%, quotes,
  backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): keep error detail in the status route's 500

test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): dismiss route rejects non-object JSON with 400

A JSON array or scalar body made `.get('alert_id')` raise, returning 500.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): let the app-wide handler answer status-route errors

CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-15 10:58:57 -04:00
committed by GitHub
co-authored by Claude Opus 5
parent d01da3bd9f
commit 869e36fb2f
25 changed files with 2938 additions and 192 deletions
+242 -182
View File
@@ -12,6 +12,8 @@ from web_interface.blueprints.api_v3 import (
get_git_version, jsonify, logger, os, request, resolve_pull_command,
shutil, subprocess,
)
import threading
import web_interface.blueprints.api_v3 as _pkg
# Read through the module rather than bound by value: tests patch these
# as module attributes, and a value binding would not see the patch.
@@ -120,6 +122,32 @@ def get_system_version():
except Exception as e:
logger.error("get_system_version failed: %s", e, exc_info=True)
return jsonify({'status': 'error', 'message': 'Unable to retrieve version'}), 500
@api_v3.route('/system/auto-update', methods=['GET'])
def get_auto_update_status():
"""Weekly automatic update status: last result, next check, and any alert.
No local except: a failure falls through to the app-wide handler in
web_interface/app.py, which logs the traceback and returns the redacted
detail -- the response every route gives, without a second copy here.
"""
from web_interface import auto_update
config = api_v3.config_manager.load_config() if api_v3.config_manager else {}
return jsonify({'status': 'success', 'data': auto_update.describe_status(config)})
@api_v3.route('/system/auto-update/dismiss', methods=['POST'])
def dismiss_auto_update_alert():
"""Hide the current automatic-update banner until a new alert replaces it."""
from web_interface import auto_update
payload = request.get_json(silent=True)
# A JSON array or scalar is a bad request, not a 500.
alert_id = str(payload.get('alert_id') or '').strip() if isinstance(payload, dict) else ''
if not alert_id:
return jsonify({'status': 'error', 'message': 'alert_id required'}), 400
auto_update.dismiss_alert(alert_id)
return jsonify({'status': 'success'})
@api_v3.route('/system/check-update', methods=['GET'])
def check_for_update():
"""Check whether a newer LEDMatrix commit is available on origin/main."""
@@ -198,6 +226,219 @@ def _sudo_hint_for(text):
return None
_core_update_lock = threading.Lock()
def perform_core_update():
"""Pull the latest LEDMatrix code and sync its dependencies.
Shared by the Overview "Update Code" button and the weekly automatic
updater (web_interface/auto_update.py), so both take exactly the same
path. Returns the JSON-able payload the button has always received:
``status``, ``message`` and ``restart_required``.
"""
# The button and the scheduler can fire together; two pulls racing
# over one checkout (and one stash) is how local changes get lost.
if not _core_update_lock.acquire(blocking=False):
return {'status': 'error', 'restart_required': False,
'message': 'An update is already in progress; try again shortly.'}
try:
return _perform_core_update_locked()
finally:
_core_update_lock.release()
def _perform_core_update_locked():
project_dir = str(PROJECT_ROOT)
# Decide how to pull BEFORE stashing. If this checkout cannot be
# updated at all, stashing first would put the user's local changes
# away for an update that was never going to run.
pull_args, upstream_note, pull_error = resolve_pull_command(project_dir)
if pull_error:
logger.warning("git pull not attempted: %s", pull_error)
return {'status': 'error', 'message': pull_error, 'restart_required': False}
# Check if there are local changes that need to be stashed
# Exclude plugins directory - plugins are separate repos and shouldn't be stashed with base project
# Use --untracked-files=no to skip untracked files check (much faster with symlinked plugins)
try:
status_result = subprocess.run(
['git', 'status', '--porcelain', '--untracked-files=no'],
capture_output=True,
text=True,
timeout=30,
cwd=project_dir
)
# Filter out any changes in plugins directory - plugins are separate repositories
# Git status format: XY filename (where X is status of index, Y is status of work tree)
status_lines = [line for line in status_result.stdout.strip().split('\n')
if line.strip() and 'plugins/' not in line]
has_changes = bool('\n'.join(status_lines).strip())
except subprocess.TimeoutExpired:
# If status check times out, assume there might be changes and proceed
# This is safer than failing the update
has_changes = True
status_result = type('obj', (object,), {'stdout': '', 'stderr': 'Status check timed out'})()
stash_info = ""
# Stash local changes if they exist (excluding plugins)
# Plugins are separate repositories and shouldn't be stashed with base project updates
if has_changes:
try:
# Use pathspec to exclude plugins directory from stash
stash_result = subprocess.run(
['git', 'stash', 'push', '-m', 'LEDMatrix auto-stash before update', '--', ':!plugins'],
capture_output=True,
text=True,
timeout=30,
cwd=project_dir
)
if stash_result.returncode == 0:
logger.debug("git stash: stashed local changes before pull")
stash_info = " Local changes were stashed."
else:
logger.warning("git stash failed before pull (returncode=%d)", stash_result.returncode)
except subprocess.TimeoutExpired:
logger.warning("git stash timed out, proceeding with pull")
# Record HEAD before the pull so dependency changes can be detected
old_head = None
try:
_pre = subprocess.run(['git', 'rev-parse', 'HEAD'],
capture_output=True, text=True, timeout=10, cwd=project_dir)
if _pre.returncode == 0:
old_head = _pre.stdout.strip()
except subprocess.TimeoutExpired:
logger.warning("git rev-parse timed out before pull")
# Whether the pull actually brought new code in. "Already up to
# date" is a success too, and prompting for a restart then would
# train users to ignore the prompt.
code_changed = False
# Requirement files whose install failed. The automatic updater refuses
# to restart onto code whose dependencies did not install.
dependency_failures = []
# Perform the git pull. Branches without an upstream were given
# an explicit "origin <branch>" above so the update still works.
result = subprocess.run(
pull_args,
capture_output=True,
text=True,
timeout=60,
cwd=project_dir
)
# Give the branch tracking information so the next pull is a plain
# `git pull` — otherwise every update repeats the fallback.
if result.returncode == 0 and upstream_note:
branch = _git_current_branch(project_dir)
if branch:
try:
subprocess.run(
['git', 'branch', f'--set-upstream-to=origin/{branch}', branch],
capture_output=True, text=True, timeout=10, cwd=project_dir)
except (subprocess.TimeoutExpired, OSError) as exc:
logger.debug("could not set upstream for %s: %s", branch, exc)
# Return custom response for git_pull
if result.returncode == 0:
pull_message = "Code updated successfully."
if has_changes:
pull_message = f"Code updated successfully. Local changes were automatically stashed.{stash_info}"
if result.stdout and "Already up to date" not in result.stdout:
pull_message = f"Code updated successfully.{stash_info}"
if upstream_note:
pull_message = f"{pull_message} {upstream_note}"
# Keep Python dependencies in sync automatically: if the pull
# changed a requirements file, install it now — users updating
# from the web UI (most of them) never SSH in to pip install.
# Installs go through the same root-visible path as the
# Tools-tab buttons (_pip_install_requirements).
dep_notes = []
try:
_post = subprocess.run(['git', 'rev-parse', 'HEAD'],
capture_output=True, text=True, timeout=10, cwd=project_dir)
new_head = _post.stdout.strip() if _post.returncode == 0 else None
if old_head and new_head and old_head != new_head:
code_changed = True
diff = subprocess.run(
['git', 'diff', '--name-only', f'{old_head}..{new_head}'],
capture_output=True, text=True, timeout=15, cwd=project_dir)
changed = set(diff.stdout.split()) if diff.returncode == 0 else set()
for rel in ('requirements.txt', 'web_interface/requirements.txt'):
req_path = PROJECT_ROOT / rel
if rel not in changed or not req_path.exists():
continue
# Each file's install is isolated: a timeout or
# OSError (e.g. the sudo wrapper/interpreter
# missing) on one file must not abort the other.
try:
r = _pip_install_requirements(req_path, timeout=180)
if r.returncode == 0:
dep_notes.append(f"Dependencies from {rel} updated.")
else:
dependency_failures.append(rel)
dep_notes.append(
f"Dependency install from {rel} failed — "
"run Install Base Requirements from the Tools tab.")
logger.warning("post-update pip install failed for %s: %s",
rel, _truncate_output(r.stdout, r.stderr))
except subprocess.TimeoutExpired:
dependency_failures.append(rel)
dep_notes.append(
f"Dependency install from {rel} timed out — "
"run Install Base Requirements from the Tools tab.")
logger.warning("post-update pip install timed out for %s", rel)
except OSError as install_err:
dependency_failures.append(rel)
dep_notes.append(
f"Dependency install from {rel} failed — "
"run Install Base Requirements from the Tools tab.")
logger.warning("post-update pip install errored for %s: %s",
rel, install_err)
except subprocess.TimeoutExpired:
logger.warning("post-update dependency sync timed out")
if dep_notes:
pull_message += " " + " ".join(dep_notes)
# A `git pull` restores built-in plugins (committed under
# plugin-repos/) even if the user uninstalled them. Re-remove
# any the user previously uninstalled so the update doesn't
# resurrect them.
if api_v3.plugin_store_manager:
try:
purged = api_v3.plugin_store_manager.purge_uninstalled_plugins()
if purged:
logger.info(
"Re-removed %d uninstalled plugin(s) restored by update: %s",
len(purged), ", ".join(purged),
)
except (OSError, RuntimeError) as purge_err:
logger.warning("Post-update plugin purge failed: %s", purge_err)
else:
logger.warning("git pull failed (returncode=%d): %s", result.returncode, result.stderr)
# Show git's own first line: "check logs" leaves the user with
# nothing to act on, and these failures are usually actionable
# (conflicting local commits, no upstream, network).
detail = next((ln.strip() for ln in (result.stderr or '').splitlines()
if ln.strip()), '')
pull_message = f"Update failed: {detail}" if detail else "Update failed; check logs for details"
# Nothing here restarts anything: the pull replaces files on
# disk while the display and web services keep running the code
# they loaded at boot. Without this the user is told the update
# succeeded and sees no change until they happen to reboot.
return {
'status': 'success' if result.returncode == 0 else 'error',
'message': pull_message,
'restart_required': bool(result.returncode == 0 and code_changed),
'dependency_failures': dependency_failures,
}
@api_v3.route('/system/action', methods=['POST'])
def execute_system_action():
"""Execute system actions (start/stop/reboot/etc)"""
@@ -265,188 +506,7 @@ def execute_system_action():
result = subprocess.run(['sudo', 'poweroff'],
capture_output=True, text=True, timeout=10)
elif action == 'git_pull':
# Use PROJECT_ROOT instead of hardcoded path
project_dir = str(PROJECT_ROOT)
# Decide how to pull BEFORE stashing. If this checkout cannot be
# updated at all, stashing first would put the user's local changes
# away for an update that was never going to run.
pull_args, upstream_note, pull_error = resolve_pull_command(project_dir)
if pull_error:
logger.warning("git pull not attempted: %s", pull_error)
return jsonify({'status': 'error', 'message': pull_error})
# Check if there are local changes that need to be stashed
# Exclude plugins directory - plugins are separate repos and shouldn't be stashed with base project
# Use --untracked-files=no to skip untracked files check (much faster with symlinked plugins)
try:
status_result = subprocess.run(
['git', 'status', '--porcelain', '--untracked-files=no'],
capture_output=True,
text=True,
timeout=30,
cwd=project_dir
)
# Filter out any changes in plugins directory - plugins are separate repositories
# Git status format: XY filename (where X is status of index, Y is status of work tree)
status_lines = [line for line in status_result.stdout.strip().split('\n')
if line.strip() and 'plugins/' not in line]
has_changes = bool('\n'.join(status_lines).strip())
except subprocess.TimeoutExpired:
# If status check times out, assume there might be changes and proceed
# This is safer than failing the update
has_changes = True
status_result = type('obj', (object,), {'stdout': '', 'stderr': 'Status check timed out'})()
stash_info = ""
# Stash local changes if they exist (excluding plugins)
# Plugins are separate repositories and shouldn't be stashed with base project updates
if has_changes:
try:
# Use pathspec to exclude plugins directory from stash
stash_result = subprocess.run(
['git', 'stash', 'push', '-m', 'LEDMatrix auto-stash before update', '--', ':!plugins'],
capture_output=True,
text=True,
timeout=30,
cwd=project_dir
)
if stash_result.returncode == 0:
logger.debug("git stash: stashed local changes before pull")
stash_info = " Local changes were stashed."
else:
logger.warning("git stash failed before pull (returncode=%d)", stash_result.returncode)
except subprocess.TimeoutExpired:
logger.warning("git stash timed out, proceeding with pull")
# Record HEAD before the pull so dependency changes can be detected
old_head = None
try:
_pre = subprocess.run(['git', 'rev-parse', 'HEAD'],
capture_output=True, text=True, timeout=10, cwd=project_dir)
if _pre.returncode == 0:
old_head = _pre.stdout.strip()
except subprocess.TimeoutExpired:
logger.warning("git rev-parse timed out before pull")
# Whether the pull actually brought new code in. "Already up to
# date" is a success too, and prompting for a restart then would
# train users to ignore the prompt.
code_changed = False
# Perform the git pull. Branches without an upstream were given
# an explicit "origin <branch>" above so the update still works.
result = subprocess.run(
pull_args,
capture_output=True,
text=True,
timeout=60,
cwd=project_dir
)
# Give the branch tracking information so the next pull is a plain
# `git pull` — otherwise every update repeats the fallback.
if result.returncode == 0 and upstream_note:
branch = _git_current_branch(project_dir)
if branch:
try:
subprocess.run(
['git', 'branch', f'--set-upstream-to=origin/{branch}', branch],
capture_output=True, text=True, timeout=10, cwd=project_dir)
except (subprocess.TimeoutExpired, OSError) as exc:
logger.debug("could not set upstream for %s: %s", branch, exc)
# Return custom response for git_pull
if result.returncode == 0:
pull_message = "Code updated successfully."
if has_changes:
pull_message = f"Code updated successfully. Local changes were automatically stashed.{stash_info}"
if result.stdout and "Already up to date" not in result.stdout:
pull_message = f"Code updated successfully.{stash_info}"
if upstream_note:
pull_message = f"{pull_message} {upstream_note}"
# Keep Python dependencies in sync automatically: if the pull
# changed a requirements file, install it now — users updating
# from the web UI (most of them) never SSH in to pip install.
# Installs go through the same root-visible path as the
# Tools-tab buttons (_pip_install_requirements).
dep_notes = []
try:
_post = subprocess.run(['git', 'rev-parse', 'HEAD'],
capture_output=True, text=True, timeout=10, cwd=project_dir)
new_head = _post.stdout.strip() if _post.returncode == 0 else None
if old_head and new_head and old_head != new_head:
code_changed = True
diff = subprocess.run(
['git', 'diff', '--name-only', f'{old_head}..{new_head}'],
capture_output=True, text=True, timeout=15, cwd=project_dir)
changed = set(diff.stdout.split()) if diff.returncode == 0 else set()
for rel in ('requirements.txt', 'web_interface/requirements.txt'):
req_path = PROJECT_ROOT / rel
if rel not in changed or not req_path.exists():
continue
# Each file's install is isolated: a timeout or
# OSError (e.g. the sudo wrapper/interpreter
# missing) on one file must not abort the other.
try:
r = _pip_install_requirements(req_path, timeout=180)
if r.returncode == 0:
dep_notes.append(f"Dependencies from {rel} updated.")
else:
dep_notes.append(
f"Dependency install from {rel} failed — "
"run Install Base Requirements from the Tools tab.")
logger.warning("post-update pip install failed for %s: %s",
rel, _truncate_output(r.stdout, r.stderr))
except subprocess.TimeoutExpired:
dep_notes.append(
f"Dependency install from {rel} timed out — "
"run Install Base Requirements from the Tools tab.")
logger.warning("post-update pip install timed out for %s", rel)
except OSError as install_err:
dep_notes.append(
f"Dependency install from {rel} failed — "
"run Install Base Requirements from the Tools tab.")
logger.warning("post-update pip install errored for %s: %s",
rel, install_err)
except subprocess.TimeoutExpired:
logger.warning("post-update dependency sync timed out")
if dep_notes:
pull_message += " " + " ".join(dep_notes)
# A `git pull` restores built-in plugins (committed under
# plugin-repos/) even if the user uninstalled them. Re-remove
# any the user previously uninstalled so the update doesn't
# resurrect them.
if api_v3.plugin_store_manager:
try:
purged = api_v3.plugin_store_manager.purge_uninstalled_plugins()
if purged:
logger.info(
"Re-removed %d uninstalled plugin(s) restored by update: %s",
len(purged), ", ".join(purged),
)
except (OSError, RuntimeError) as purge_err:
logger.warning("Post-update plugin purge failed: %s", purge_err)
else:
logger.warning("git pull failed (returncode=%d): %s", result.returncode, result.stderr)
# Show git's own first line: "check logs" leaves the user with
# nothing to act on, and these failures are usually actionable
# (conflicting local commits, no upstream, network).
detail = next((ln.strip() for ln in (result.stderr or '').splitlines()
if ln.strip()), '')
pull_message = f"Update failed: {detail}" if detail else "Update failed; check logs for details"
# Nothing here restarts anything: the pull replaces files on
# disk while the display and web services keep running the code
# they loaded at boot. Without this the user is told the update
# succeeded and sees no change until they happen to reboot.
return jsonify({
'status': 'success' if result.returncode == 0 else 'error',
'message': pull_message,
'restart_required': bool(result.returncode == 0 and code_changed),
})
return jsonify(perform_core_update())
elif action == 'checkout_branch':
# Switch branches from the Tools tab. Needed because a checkout
# that predates tracking (or a restored backup) can leave the pi