mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 22:35:08 +00:00
* refactor(api-v3): split the 10,469-line blueprint into a package web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181 functions in one module -- 9% of the core by line count and three times the next largest file. It becomes a package of nine route modules grouped by path segment, plus __init__.py for the shared imports, constants, Blueprint and helpers. Every route module decorates the SAME api_v3 Blueprint object, so endpoint names stay api_v3.<function>, the URL map is unchanged and app.py is untouched. Verified: 111 routes before, 111 after, byte-identical rules, endpoints and methods, and every endpoint still on the one blueprint. plugins 3,867 config 1,178 starlark 692 system 619 fonts 452 misc 398 wifi 361 display 326 backup 212 __init__ 1,787 (imports, constants, Blueprint, 56 helpers) Two things the URL-map check could not catch, both found by running the suite: 1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one directory deeper made that resolve to web_interface/ instead of the project root. Nothing failed at import; it surfaced as ~110 tests failing with 404s and "installation script not found", because every path built from it was one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts PROJECT_ROOT/run.py exists so the next move cannot repeat it. 2. Module-attribute patching. Tests do monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route module that binds such a name by value never sees the patch. The shared code therefore stays in __init__.py rather than moving to a _common submodule -- it has to live on the module the tests patch -- and the eleven names tests patch are read back through the package (_pkg.X) instead of bound by value. Those eleven were found by AST-scanning every setattr in the test tree, not by guessing; "time" is among them, used to drive a fake clock through the second-resolution credential-backup filenames. Test changes are confined to what genuinely moved: patch targets that now name the owning route module, imports of helpers, and six tests that scan the api_v3 source as a file and now read the package directory. Full suite: 4,278 passed, 68 skipped, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(api-v3): address CodeRabbit findings from the blueprint-split review Fixes to the api_v3 package split (PR #553), one per finding verified against the actual code: - __init__.py: _redact_credentials only blanked scalar values under a credential-named key; a bare list of secrets under such a key (e.g. tokens: ["a", "b"]) passed through untouched, since the list branch recursed with no memory that its key looked like a credential. Nested dicts still walk normally (a documented, tested behaviour -- a container like secrets: {api_key: ..., note: ...} is a section name, not a value to blank outright), but any value reached under a credential-shaped key is now actually blanked. - __init__.py: the OAuth helper script's raw stderr/stdout went to logger.error unredacted (CWE-532) right next to a comment claiming this was deliberate; the HTTP response already used the existing redact_text helper. Routed the log line through the same helper. - __init__.py / starlark.py: the standalone Starlark manifest fallback (used when the plugin instance isn't loaded) read-modified-wrote manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe (plugin-repos/starlark-apps/manager.py), which already holds an flock for the same file when the plugin is loaded. Added _starlark_manifest_lock, mirroring that pattern, and wrapped every standalone read-modify-write call site in it. The app-config update route also wrote config.json and the manifest as two separate, non-transactional writes (a second, distinct finding at the same call site); config.json is now rolled back if the manifest write that follows it fails. - backup.py: restore options used bare bool() on values from the request, so {"restore_secrets": "false"} restored secrets anyway (bool("false") is True). Switched to the existing _coerce_to_bool helper already used for this exact purpose elsewhere in the package. - config.py: an automated import-rewrite mangled four user-facing validation strings and their neighbouring comments -- "Invalid start time" had become "Invalid start _pkg.time" (and likewise for "end time") in both the schedule and dim-schedule per-day validation paths. - display.py: `import _pkg.time as time_module` -- _pkg is a local alias for the package, not a real importable module, so this raised ModuleNotFoundError whenever a caller restarted an already-running display service via /display/on-demand/start, after the on-demand request was already written to cache. Fixed to `import time`. Audited the rest of the package for the same `_pkg.<module>` import mistake; every other `_pkg.` reference is a legitimate attribute read-through (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken import statement. - fonts.py: validate_file_upload's max_size_mb parameter is silently unused by that helper (it only checks filename/extension) -- the font upload route saved arbitrarily large files as a result. Added the same seek-and-check pattern already used for the sibling .star upload. - wifi.py: two ad hoc, inconsistent bool coercions. POST /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled it. POST /wifi/radio's enabled/force parsing recognized real bool and some strings but not int 1/0 (1 is True is False in Python). Factored one small _parse_bool_ish helper local to this file and used it at all three sites. Not changed: the "unknown/misspelled restore option keys default to True" half of the backup.py finding -- the file's own comment documents that a missing key deliberately means "restore everything," matching the already-existing JSON-parse-failure guard a few lines above it; only the bool-coercion defect was a real bug. Added or extended regression tests for every fix, following each area's existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed on both this branch and origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change) -- no new failures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5 * fix(api-v3): reject unknown restore option keys CodeRabbit's review of the blueprint split (#553) asked that POST /backup/restore reject option keys outside RestoreOptions' known set. The follow-up commit fixed the bool("false")-is-True bug with _coerce_to_bool but never added the key check: a typo'd or renamed key (e.g. "restoreSecrets") is silently ignored by opts_dict.get(key, True), so the flag stays at its True default and secrets get restored despite the caller's request saying otherwise -- with no indication anything was wrong. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb * fix(api-v3): address CodeRabbit findings on the blueprint split - _redact_credentials: blank scalar descendants of objects reached through a credential-owned list (e.g. tokens: [{"value": "secret"}]) regardless of field name -- the existing name-based walk only protected direct dict values under a credential key, not list items. - wifi.py: reject enabled/force/auto_enable_ap_mode values _parse_bool_ish can't recognize (400) instead of silently treating them as False, which could disable Wi-Fi or the radio itself. - Starlark manifest locking: lock a stable manifest.json.lock sidecar instead of manifest.json itself, in both the standalone route path (_starlark_manifest_lock) and the plugin path (StarlarkAppsPlugin._save_manifest / _update_manifest_safe). manifest.json is replaced by an atomic rename on every write, which swaps in a fresh inode; a lock held on the old inode does not exclude a second locker that opens the path afresh right after the rename and gets the new inode, so two writers could race despite each holding "a lock". A sidecar that no write ever touches always resolves to the same inode for every locker. Skipped as stale: the "serialize the complete manifest read-modify-write" finding at api_v3/__init__.py -- every standalone handler that calls _write_starlark_manifest is already wrapped in _starlark_manifest_lock() on this branch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): re-check reconciliation findings by the reconciler's own rules Both CodeRabbit findings on the merge commit, verified against the code first. Major, plugins.py: the stale-findings filter derived its own notion of "in config" and "on disk", and both were looser than the reconciliation module's. set(load_config()) also contains system keys, the secrets-file keys load_config() merges in, and non-dict values; and any directory holding a manifest.json counted as installed even when that manifest does not parse. Either looseness clears a finding that is still true -- and a secrets key read as a plugin is the precise bug the filter exists to stop reporting, so reintroducing that asymmetry while re-checking was the wrong way round. The two extractions now live in state_reconciliation.py as config_plugin_ids() and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys() alongside. _get_config_state() and _get_disk_state() use them too, so there is one definition rather than two that can drift. _get_disk_state() re-reads each manifest for version/name after taking membership from the shared extractor; that costs one extra small read per plugin on a path that runs once per boot. Minor, the new test: the fixture assigned api_v3.config_manager and api_v3.plugin_manager directly. Those live on a module-level blueprint singleton, so the mocks leaked into every later test that imports api_v3 -- pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr, which restores them. This is the same pollution class that made an earlier test in this session break seven unrelated ones, so it is worth getting right. Five cases added for the parity itself: a secrets key, a system key and a non-dict value must not clear an "installed but missing from config" finding, and neither an unparseable manifest nor a .standalone-backup- directory may count as installed. All five fail against the looser version. Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and CodeRabbit all pass. Codacy reads action_required on every commit of this branch including the first, so it is pre-existing and not from this work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
303 lines
13 KiB
Python
303 lines
13 KiB
Python
"""Tests for surfacing the underlying error in web responses.
|
|
|
|
Regression under test: every failing endpoint returned "An error occurred; see
|
|
logs for details" and nothing else. On a device whose storage was failing that
|
|
sentence came back from the restart action, from /system/status, and from
|
|
/logs -- the log viewer itself -- because journalctl could not be executed. The
|
|
exception underneath said `[Errno 5] Input/output error: 'systemctl'`, which
|
|
names the fault outright, and nine handlers were discarding it entirely rather
|
|
than even logging it.
|
|
"""
|
|
|
|
import pytest
|
|
|
|
from src.web_interface.error_handler import describe_exception
|
|
|
|
|
|
class TestDescribeException:
|
|
def test_names_the_type_and_message(self):
|
|
detail = describe_exception(OSError(5, "Input/output error", "systemctl"))
|
|
assert detail == "OSError: [Errno 5] Input/output error: 'systemctl'"
|
|
|
|
def test_the_reported_failure_is_legible(self):
|
|
# The whole point: this string is the diagnosis.
|
|
assert "Input/output error" in describe_exception(
|
|
OSError(5, "Input/output error", "systemctl"))
|
|
|
|
def test_a_bare_exception_still_names_its_type(self):
|
|
# A PermissionError with no message still says more than "unknown".
|
|
assert describe_exception(PermissionError()) == "PermissionError"
|
|
assert describe_exception(Exception()) == "Exception"
|
|
|
|
def test_message_is_kept_when_present(self):
|
|
assert describe_exception(ValueError("bad port")) == "ValueError: bad port"
|
|
|
|
|
|
class TestCredentialRedaction:
|
|
"""Exception text quotes URLs, and plugins authenticate by query string."""
|
|
|
|
@pytest.mark.parametrize("secret_text,leaked", [
|
|
("failed: https://api.x.com/v1?api_key=SEC123&city=Tampa", "SEC123"),
|
|
("token=abcdef123456 was rejected", "abcdef123456"),
|
|
("connect failed password=hunter2", "hunter2"),
|
|
("GET /?access_token=zzz999", "zzz999"),
|
|
('{"secret": "topsecret"}', "topsecret"),
|
|
# requests quotes the URL it failed on, and both of these forms turn
|
|
# up in real client exceptions.
|
|
("401 for https://user:hunter2@example.com/api", "hunter2"),
|
|
("headers: {'Authorization': 'Bearer eyJ.SECRET.sig'}", "eyJ.SECRET.sig"),
|
|
("Authorization: Basic dXNlcjpwYXNzd29yZA==", "dXNlcjpwYXNzd29yZA=="),
|
|
("Proxy-Authorization: Bearer ptok999", "ptok999"),
|
|
# Any scheme, not a fixed list -- a list silently leaks whatever it
|
|
# does not name, and plugin APIs invent their own.
|
|
("Authorization: ApiKey SECRET123", "SECRET123"),
|
|
("Authorization: Negotiate YIIZnegotiateblob", "YIIZnegotiateblob"),
|
|
("Authorization: NTLM TlRMTVNTUAAB", "TlRMTVNTUAAB"),
|
|
("authorization: barecredential", "barecredential"),
|
|
])
|
|
def test_credentials_never_reach_the_response(self, secret_text, leaked):
|
|
detail = describe_exception(RuntimeError(secret_text))
|
|
assert leaked not in detail
|
|
assert "<redacted>" in detail
|
|
|
|
def test_the_parameter_name_survives_redaction(self):
|
|
# Knowing *which* credential was involved is part of the diagnosis.
|
|
detail = describe_exception(RuntimeError("https://x/y?api_key=SEC123"))
|
|
assert "api_key" in detail
|
|
|
|
def test_unknown_schemes_keep_their_name(self):
|
|
for scheme in ("ApiKey", "Negotiate", "NTLM", "AWS4-HMAC-SHA256"):
|
|
detail = describe_exception(
|
|
RuntimeError("Authorization: %s SECRETVALUE" % scheme))
|
|
assert scheme in detail, detail
|
|
assert "SECRETVALUE" not in detail, detail
|
|
|
|
def test_auth_scheme_and_username_survive(self):
|
|
# Which kind of credential, and whose, without the credential itself.
|
|
assert "Bearer" in describe_exception(
|
|
RuntimeError("Authorization: Bearer eyJ.SECRET.sig"))
|
|
assert "user" in describe_exception(
|
|
RuntimeError("https://user:hunter2@example.com"))
|
|
|
|
def test_non_secret_context_is_preserved(self):
|
|
detail = describe_exception(RuntimeError("https://api.x.com/v1?city=Tampa"))
|
|
assert "city=Tampa" in detail
|
|
assert "<redacted>" not in detail
|
|
|
|
|
|
class TestBounds:
|
|
def test_long_messages_are_truncated(self):
|
|
detail = describe_exception(ValueError("x" * 5000))
|
|
assert len(detail) <= 400
|
|
|
|
def test_newlines_are_collapsed_to_one_line(self):
|
|
detail = describe_exception(ValueError("line one\nline two\tthree"))
|
|
assert "\n" not in detail and "\t" not in detail
|
|
assert detail == "ValueError: line one line two three"
|
|
|
|
def test_custom_length_is_honoured(self):
|
|
assert len(describe_exception(ValueError("y" * 500), max_length=50)) <= 50
|
|
|
|
|
|
class TestHandlersCarryDetail:
|
|
"""The response shape callers actually see."""
|
|
|
|
def test_no_api_v3_handler_discards_its_exception(self):
|
|
"""Every generic-message handler must log a traceback and return detail.
|
|
|
|
Nine of them bound `e` and never used it, so the promised log entry was
|
|
never written either. Checking merely that *something* was logged is
|
|
too weak -- a `logger.info("failed")` would satisfy it while throwing
|
|
the exception away just as completely, so this asserts the two things
|
|
that actually make the failure diagnosable: an error-level record with
|
|
the traceback, and the sanitized detail in the response.
|
|
"""
|
|
import ast
|
|
|
|
# api_v3 is a package; the routes are spread across its modules.
|
|
import pathlib
|
|
src = "\n".join(
|
|
p.read_text() for p in
|
|
sorted(pathlib.Path("web_interface/blueprints/api_v3").glob("*.py")))
|
|
tree = ast.parse(src)
|
|
|
|
# This used to match one exact message string, so a handler that wrote
|
|
# its own wording was never checked. All thirteen Starlark routes did
|
|
# -- "Failed to browse repository" and friends -- and every one of them
|
|
# answered a 500 with no detail at all, which is how the app store
|
|
# spent three releases failing for reasons nobody could read. The rule
|
|
# is now the shape that matters: if it returns 5xx, it says why.
|
|
PRE_EXISTING = {
|
|
# Not part of this change. This set may shrink, never grow.
|
|
'backup_delete', 'backup_export', 'backup_list', 'backup_preview',
|
|
'backup_restore', 'backup_validate', 'checkout_branch',
|
|
'execute_system_action', 'get_git_branches', 'get_git_info',
|
|
'get_hardware_status', 'get_logs', 'get_system_status',
|
|
'get_system_version', 'scan_wifi_networks',
|
|
}
|
|
|
|
def enclosing_function(handler):
|
|
"""Innermost function containing `handler`."""
|
|
best = None
|
|
for fn in [n for n in ast.walk(tree)
|
|
if isinstance(n, (ast.FunctionDef, ast.AsyncFunctionDef))]:
|
|
if any(h is handler for h in ast.walk(fn)):
|
|
if best is None or fn.lineno > best.lineno:
|
|
best = fn
|
|
return best.name if best else '<module>'
|
|
|
|
def only_catches_importerror(handler):
|
|
"""An `except ImportError` arm and nothing else.
|
|
|
|
A missing optional dependency is a configuration fact, not a
|
|
crash: the module name is the whole diagnosis and it is already
|
|
in the response, so a stack trace would be noise. Detail is still
|
|
required -- only the traceback log is excused.
|
|
"""
|
|
t = handler.type
|
|
names = ([t] if isinstance(t, ast.Name)
|
|
else list(t.elts) if isinstance(t, ast.Tuple) else [])
|
|
return bool(names) and all(
|
|
isinstance(n, ast.Name) and n.id == 'ImportError' for n in names)
|
|
|
|
def returns_5xx(handler):
|
|
for r in [n for n in ast.walk(handler) if isinstance(n, ast.Return)]:
|
|
v = r.value
|
|
if isinstance(v, ast.Tuple) and len(v.elts) == 2:
|
|
code = v.elts[1]
|
|
if (isinstance(code, ast.Constant)
|
|
and isinstance(code.value, int)
|
|
and 500 <= code.value < 600):
|
|
return True
|
|
return False
|
|
|
|
def logs_a_traceback(handler):
|
|
"""An error/exception-level log call carrying exc_info."""
|
|
for call in [n for n in ast.walk(handler) if isinstance(n, ast.Call)]:
|
|
func = call.func
|
|
if not isinstance(func, ast.Attribute):
|
|
continue
|
|
if func.attr == "exception": # implies exc_info
|
|
return True
|
|
if func.attr not in ("error", "critical"):
|
|
continue
|
|
if any(kw.arg == "exc_info" and getattr(kw.value, "value", False) is True
|
|
for kw in call.keywords):
|
|
return True
|
|
return False
|
|
|
|
def describes_this_exception(node, bound):
|
|
"""A describe_exception(<bound>) call anywhere under `node`."""
|
|
for call in [n for n in ast.walk(node) if isinstance(n, ast.Call)]:
|
|
if not (isinstance(call.func, ast.Name)
|
|
and call.func.id == "describe_exception"):
|
|
continue
|
|
if bound is None:
|
|
return True # bare `except:` cannot name it; accept
|
|
if any(isinstance(a, ast.Name) and a.id == bound
|
|
for a in call.args):
|
|
return True
|
|
return False
|
|
|
|
def returns_the_detail(handler):
|
|
"""The detail must be inside what the handler actually returns.
|
|
|
|
Looking anywhere in the handler is too weak: a handler could
|
|
compute describe_exception(e), drop it on the floor, and return the
|
|
generic message with no details field, while still passing. So the
|
|
call has to appear within a `return` expression.
|
|
"""
|
|
returns = [n for n in ast.walk(handler) if isinstance(n, ast.Return)]
|
|
if not returns:
|
|
return False
|
|
return all(describes_this_exception(r, handler.name) for r in returns)
|
|
|
|
offenders = []
|
|
for h in [n for n in ast.walk(tree) if isinstance(n, ast.ExceptHandler)]:
|
|
if not returns_5xx(h):
|
|
continue
|
|
if enclosing_function(h) in PRE_EXISTING:
|
|
continue
|
|
missing = []
|
|
if not logs_a_traceback(h) and not only_catches_importerror(h):
|
|
missing.append("error-level log with exc_info")
|
|
if not returns_the_detail(h):
|
|
missing.append("describe_exception(e) in the response")
|
|
if missing:
|
|
offenders.append((h.lineno, missing))
|
|
|
|
assert not offenders, (
|
|
"handlers returning the generic message without %s: %r"
|
|
% ("both a traceback log and the detail", offenders))
|
|
|
|
def test_client_errors_keep_their_own_status(self):
|
|
"""A 405 must not be reported as a server-side UNKNOWN_ERROR.
|
|
|
|
Werkzeug's HTTPExceptions subclass Exception, so the catch-all saw them
|
|
too: a GET on a POST-only route came back 500 "an error occurred",
|
|
which tells the caller nothing and blames the wrong side. Found while
|
|
probing a device whose POST-only config endpoints answered every GET
|
|
with UNKNOWN_ERROR.
|
|
"""
|
|
from flask import Flask, jsonify
|
|
from werkzeug.exceptions import HTTPException
|
|
|
|
app = Flask(__name__)
|
|
|
|
@app.errorhandler(Exception)
|
|
def handle(error):
|
|
if isinstance(error, HTTPException):
|
|
return jsonify({
|
|
"status": "error",
|
|
"error_code": (error.name or "HTTP_ERROR").upper().replace(" ", "_"),
|
|
"message": error.description,
|
|
}), error.code or 500
|
|
return jsonify({
|
|
"status": "error",
|
|
"error_code": "UNKNOWN_ERROR",
|
|
"message": "An error occurred; see logs for details",
|
|
"details": describe_exception(error),
|
|
}), 500
|
|
|
|
@app.route("/only-post", methods=["POST"])
|
|
def only_post():
|
|
return jsonify({"ok": True})
|
|
|
|
@app.route("/boom")
|
|
def boom():
|
|
raise OSError(5, "Input/output error", "systemctl")
|
|
|
|
client = app.test_client()
|
|
|
|
resp = client.get("/only-post")
|
|
assert resp.status_code == 405, "a wrong method must stay a 405"
|
|
assert resp.get_json()["error_code"] == "METHOD_NOT_ALLOWED"
|
|
|
|
# A genuine server fault still reports as one, with its detail.
|
|
resp = client.get("/boom")
|
|
assert resp.status_code == 500
|
|
assert "Input/output error" in resp.get_json()["details"]
|
|
|
|
def test_global_handler_reports_the_underlying_error(self):
|
|
from flask import Flask, jsonify
|
|
|
|
app = Flask(__name__)
|
|
|
|
@app.errorhandler(Exception)
|
|
def handle(error):
|
|
return jsonify({
|
|
"status": "error",
|
|
"error_code": "UNKNOWN_ERROR",
|
|
"message": "An error occurred; see logs for details",
|
|
"details": describe_exception(error),
|
|
}), 500
|
|
|
|
@app.route("/boom")
|
|
def boom():
|
|
raise OSError(5, "Input/output error", "systemctl")
|
|
|
|
client = app.test_client()
|
|
body = client.get("/boom").get_json()
|
|
assert body["error_code"] == "UNKNOWN_ERROR"
|
|
assert "Input/output error" in body["details"]
|