mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
dac71afedcbc59b9aebfbbb4c910bb4e61ee33f0
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bdb9a94033 |
refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181 functions in one module -- 9% of the core by line count and three times the next largest file. It becomes a package of nine route modules grouped by path segment, plus __init__.py for the shared imports, constants, Blueprint and helpers. Every route module decorates the SAME api_v3 Blueprint object, so endpoint names stay api_v3.<function>, the URL map is unchanged and app.py is untouched. Verified: 111 routes before, 111 after, byte-identical rules, endpoints and methods, and every endpoint still on the one blueprint. plugins 3,867 config 1,178 starlark 692 system 619 fonts 452 misc 398 wifi 361 display 326 backup 212 __init__ 1,787 (imports, constants, Blueprint, 56 helpers) Two things the URL-map check could not catch, both found by running the suite: 1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one directory deeper made that resolve to web_interface/ instead of the project root. Nothing failed at import; it surfaced as ~110 tests failing with 404s and "installation script not found", because every path built from it was one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts PROJECT_ROOT/run.py exists so the next move cannot repeat it. 2. Module-attribute patching. Tests do monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route module that binds such a name by value never sees the patch. The shared code therefore stays in __init__.py rather than moving to a _common submodule -- it has to live on the module the tests patch -- and the eleven names tests patch are read back through the package (_pkg.X) instead of bound by value. Those eleven were found by AST-scanning every setattr in the test tree, not by guessing; "time" is among them, used to drive a fake clock through the second-resolution credential-backup filenames. Test changes are confined to what genuinely moved: patch targets that now name the owning route module, imports of helpers, and six tests that scan the api_v3 source as a file and now read the package directory. Full suite: 4,278 passed, 68 skipped, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(api-v3): address CodeRabbit findings from the blueprint-split review Fixes to the api_v3 package split (PR #553), one per finding verified against the actual code: - __init__.py: _redact_credentials only blanked scalar values under a credential-named key; a bare list of secrets under such a key (e.g. tokens: ["a", "b"]) passed through untouched, since the list branch recursed with no memory that its key looked like a credential. Nested dicts still walk normally (a documented, tested behaviour -- a container like secrets: {api_key: ..., note: ...} is a section name, not a value to blank outright), but any value reached under a credential-shaped key is now actually blanked. - __init__.py: the OAuth helper script's raw stderr/stdout went to logger.error unredacted (CWE-532) right next to a comment claiming this was deliberate; the HTTP response already used the existing redact_text helper. Routed the log line through the same helper. - __init__.py / starlark.py: the standalone Starlark manifest fallback (used when the plugin instance isn't loaded) read-modified-wrote manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe (plugin-repos/starlark-apps/manager.py), which already holds an flock for the same file when the plugin is loaded. Added _starlark_manifest_lock, mirroring that pattern, and wrapped every standalone read-modify-write call site in it. The app-config update route also wrote config.json and the manifest as two separate, non-transactional writes (a second, distinct finding at the same call site); config.json is now rolled back if the manifest write that follows it fails. - backup.py: restore options used bare bool() on values from the request, so {"restore_secrets": "false"} restored secrets anyway (bool("false") is True). Switched to the existing _coerce_to_bool helper already used for this exact purpose elsewhere in the package. - config.py: an automated import-rewrite mangled four user-facing validation strings and their neighbouring comments -- "Invalid start time" had become "Invalid start _pkg.time" (and likewise for "end time") in both the schedule and dim-schedule per-day validation paths. - display.py: `import _pkg.time as time_module` -- _pkg is a local alias for the package, not a real importable module, so this raised ModuleNotFoundError whenever a caller restarted an already-running display service via /display/on-demand/start, after the on-demand request was already written to cache. Fixed to `import time`. Audited the rest of the package for the same `_pkg.<module>` import mistake; every other `_pkg.` reference is a legitimate attribute read-through (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken import statement. - fonts.py: validate_file_upload's max_size_mb parameter is silently unused by that helper (it only checks filename/extension) -- the font upload route saved arbitrarily large files as a result. Added the same seek-and-check pattern already used for the sibling .star upload. - wifi.py: two ad hoc, inconsistent bool coercions. POST /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled it. POST /wifi/radio's enabled/force parsing recognized real bool and some strings but not int 1/0 (1 is True is False in Python). Factored one small _parse_bool_ish helper local to this file and used it at all three sites. Not changed: the "unknown/misspelled restore option keys default to True" half of the backup.py finding -- the file's own comment documents that a missing key deliberately means "restore everything," matching the already-existing JSON-parse-failure guard a few lines above it; only the bool-coercion defect was a real bug. Added or extended regression tests for every fix, following each area's existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed on both this branch and origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change) -- no new failures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5 * fix(api-v3): reject unknown restore option keys CodeRabbit's review of the blueprint split (#553) asked that POST /backup/restore reject option keys outside RestoreOptions' known set. The follow-up commit fixed the bool("false")-is-True bug with _coerce_to_bool but never added the key check: a typo'd or renamed key (e.g. "restoreSecrets") is silently ignored by opts_dict.get(key, True), so the flag stays at its True default and secrets get restored despite the caller's request saying otherwise -- with no indication anything was wrong. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb * fix(api-v3): address CodeRabbit findings on the blueprint split - _redact_credentials: blank scalar descendants of objects reached through a credential-owned list (e.g. tokens: [{"value": "secret"}]) regardless of field name -- the existing name-based walk only protected direct dict values under a credential key, not list items. - wifi.py: reject enabled/force/auto_enable_ap_mode values _parse_bool_ish can't recognize (400) instead of silently treating them as False, which could disable Wi-Fi or the radio itself. - Starlark manifest locking: lock a stable manifest.json.lock sidecar instead of manifest.json itself, in both the standalone route path (_starlark_manifest_lock) and the plugin path (StarlarkAppsPlugin._save_manifest / _update_manifest_safe). manifest.json is replaced by an atomic rename on every write, which swaps in a fresh inode; a lock held on the old inode does not exclude a second locker that opens the path afresh right after the rename and gets the new inode, so two writers could race despite each holding "a lock". A sidecar that no write ever touches always resolves to the same inode for every locker. Skipped as stale: the "serialize the complete manifest read-modify-write" finding at api_v3/__init__.py -- every standalone handler that calls _write_starlark_manifest is already wrapped in _starlark_manifest_lock() on this branch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): re-check reconciliation findings by the reconciler's own rules Both CodeRabbit findings on the merge commit, verified against the code first. Major, plugins.py: the stale-findings filter derived its own notion of "in config" and "on disk", and both were looser than the reconciliation module's. set(load_config()) also contains system keys, the secrets-file keys load_config() merges in, and non-dict values; and any directory holding a manifest.json counted as installed even when that manifest does not parse. Either looseness clears a finding that is still true -- and a secrets key read as a plugin is the precise bug the filter exists to stop reporting, so reintroducing that asymmetry while re-checking was the wrong way round. The two extractions now live in state_reconciliation.py as config_plugin_ids() and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys() alongside. _get_config_state() and _get_disk_state() use them too, so there is one definition rather than two that can drift. _get_disk_state() re-reads each manifest for version/name after taking membership from the shared extractor; that costs one extra small read per plugin on a path that runs once per boot. Minor, the new test: the fixture assigned api_v3.config_manager and api_v3.plugin_manager directly. Those live on a module-level blueprint singleton, so the mocks leaked into every later test that imports api_v3 -- pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr, which restores them. This is the same pollution class that made an earlier test in this session break seven unrelated ones, so it is worth getting right. Five cases added for the parity itself: a secrets key, a system key and a non-dict value must not clear an "installed but missing from config" finding, and neither an unparseable manifest nor a .standalone-backup- directory may count as installed. All five fail against the looser version. Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and CodeRabbit all pass. Codacy reads action_required on every commit of this branch including the first, so it is pre-existing and not from this work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26769ee37f |
fix(starlark): the store authenticated with a key nothing writes (#541)
#535 restored the thirteen routes, so the store stopped answering 404 -- and still would not load. Confirmed against a running device before anything was changed: /repository/browse answers 200 with 1000 apps in 27s, so the routes are fine. Two things underneath them are not. **The store never used the token the user configured.** The three repository routes read `github_token` off config.json. Nothing writes that key -- it is not in config.template.json, no setting offers it, and it appears nowhere else in the codebase. The configured token goes to config_secrets.json as `github.api_token`, which PluginStoreManager loads and every other GitHub caller uses. So the store could never be authenticated: 60 requests/hour, on the same per-IP budget 48 installed plugins spend on update checks, while the 5000 the user had already configured sat unused. On the device, /plugins/store/github-status reported authenticated with a limit of 5000 at the same moment /starlark/repository/browse reported 60, with 18 left. The store going blank was that 60 running out. **Every failure looked identical.** list_all_apps_cached turned any listing failure -- rate limit, DNS, timeout, non-200 -- into an empty app list, and the route sent that out as `status: success`, so a rate limit and an empty repository drew the same blank grid with no error anywhere. It now returns the reason, the route answers 502 with it, and a failure is no longer cached as an empty repository for two hours. The guard for a bad response was itself a crash: _make_request catches `(json.JSONDecodeError, ValueError)` but `json` was never imported, so evaluating the tuple raises NameError and the guard written for exactly this case never ran. Reachable whenever something on the path answers with HTML -- a captive portal, a proxy page, a DNS-hijacking router. Seventeen handlers answered 5xx with no detail at all. test_no_api_v3_handler_discards_its_exception is meant to prevent that across api_v3, but it matched one exact message string, and all thirteen Starlark routes wrote their own wording. The guard now keys on the shape that matters: if it returns 5xx, it says why. The 15 pre-existing non-Starlark functions are listed as a set that may shrink, never grow. **The listing was capped at 1000 and did not say so.** The contents API truncates a directory silently; tronbyt/apps has 1075 app directories, so the store showed a truncated repository and looked complete doing it. Now listed via the git trees API, which reports `truncated`, with the contents API kept as a fallback. Not addressed: the 27-second cold load -- 1075 manifests fetched five at a time behind skeleton placeholders -- which is probably the largest part of what "does not load" feels like, and wants its own change. 25 new tests. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
44f59ede07 |
fix(web): say what actually went wrong instead of "unknown" (#448)
* fix(web): say what actually went wrong instead of "unknown"
Every failing endpoint returned "An error occurred; see logs for
details" and nothing else. That is survivable until the logs are the
thing you cannot reach: a device whose SD card was failing answered the
restart action, /system/status and /logs with that same sentence -- the
log viewer included, because journalctl could not be executed -- while
the exception underneath said
[Errno 5] Input/output error: 'systemctl'
which names the fault outright. The only endpoint that helped was
/health, and only because it happens to pass a subprocess's stderr
through. Diagnosis came down to guessing which endpoint leaked something.
Add describe_exception(), returning "TypeName: message" on one line, and
populate the `details` field that the response schema has always had and
nothing ever filled. The type alone carries information -- a bare
PermissionError says more than any generic sentence.
Exception text is not automatically safe to echo: a requests error
quotes the URL it failed on, and plugins that authenticate by query
string put their key there. Credential values are redacted while the
parameter name is kept, since knowing which credential was involved is
part of the diagnosis. Length is capped and newlines collapsed so a
parser's context cannot flood a JSON field.
Nine handlers in api_v3 bound the exception and never used it, so the
promised log entry was never written either -- "see logs for details"
was false, not merely unhelpful. Those now log with a traceback and
carry the detail. The other 60 already logged and are unchanged; they
can adopt the helper as they are touched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): redact auth headers and URL userinfo, and cover every handler
Three review findings.
The sanitizer missed two credential shapes that requests puts in its
exception text verbatim: `Authorization: Bearer <token>` and
`https://user:password@host`. Both would have gone straight into a
response. The auth-scheme name and the username are kept -- they say
which credential and whose without being the secret.
The AST test only asked whether *something* had been logged, so a
`logger.info("failed")` satisfied it while discarding the exception just
as completely. It now requires an error-level record carrying exc_info
and `describe_exception()` called on the handler's own bound exception.
Enforcing that revealed the first cut had scoped itself wrongly. I had
converted the nine handlers that logged nothing and left the sixty that
logged, reasoning their detail was at least in the journal. But
/system/status is one of the sixty, and on the failing device it told me
nothing -- the journal was exactly what could not be read. Splitting
them left most of the diagnostic surface unhelpful for the case this
change exists for, so all sixty-nine now carry the detail.
Two handlers had no bound exception name, and three passed the message
through a variable rather than a literal; both shapes needed doing by
hand. Full suite: 2383 passed, one pre-existing unrelated failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): stop reporting client errors as server faults
Werkzeug's HTTPExceptions subclass Exception, so the catch-all handler
saw them too and turned every 405, 400, 413 and 415 into a 500
UNKNOWN_ERROR. A GET on a POST-only route answered "an error occurred;
see logs for details", which tells the caller nothing and blames the
wrong side -- found while probing a device whose POST-only config
endpoints did exactly that.
Hand HTTPExceptions back as themselves, with their own status and
description. A genuine server fault still reports as one, with the
detail this branch adds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): redact any auth scheme, and require the detail in the response
Two review findings.
The auth-header pattern listed Bearer, Basic, Digest and Token, so
`Authorization: ApiKey SECRET` or `Negotiate SECRET` went to the client
intact. A fixed list silently leaks whatever it does not name, and
plugin APIs invent their own schemes, so match any scheme name and keep
it while redacting the credential.
The AST test accepted a describe_exception(e) call anywhere in the
handler, which a handler could satisfy by computing the detail and
dropping it before returning the generic message. It now requires the
call inside every return expression, which is where it has to be to
reach the caller.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|