fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)

* fix(errors): serve /api/v3/errors/* from the display service's aggregator

The error aggregator is a per-process singleton and only the display
service runs plugins, so only its aggregator records anything. The routes
read the web process's own, empty one and always reported no errors.

The display service now publishes a bounded snapshot of its aggregator to
the shared cache (plugin_error_snapshot) from a daemon thread: at most once
every 10 s and only when something changed, never raising into the caller.
The routes read it and keep their response shapes, adding
snapshot_available, generated_at and clear_pending; exception text has
credentials redacted.

POST /errors/clear writes a clear request (plugin_error_clear_request) that
the display applies on its next 5 s tick via the new clear_before(), which
keeps errors recorded after the cutoff and rebuilds the counts. Until the
snapshot acknowledges the request, reads hide everything before the cutoff,
so a snapshot written just before the click cannot bring errors back. Adds
"all": true; cleared_count is null when only the display can know it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): show plugin errors in the Logs tab

A compact panel under the log viewer: per-plugin error counts, repeating
errors (type, count, affected plugins, a sample message, last seen) and a
Clear button, with empty states for "no errors" and "display service
hasn't reported yet". Polls every 15 s while the tab is active; all text
goes through escapeHtml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe where plugin error reports come from and how clear works

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(errors): redact the published snapshot before clipping it

Keeping only a traceback's tail (or clipping a message) could cut an
`api_key=` marker off while keeping the secret after it, and the web side's
redaction would then have nothing to match. The display now redacts every
free-text field of the snapshot first. The patterns move to a Flask-free
src/redaction.py so the display service can use them; redact_text in the web
error handler uses the same function, unchanged in behaviour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-23 14:32:02 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent cd5a4e2251
commit 4a1fd7464a
10 changed files with 1327 additions and 79 deletions
+79 -5
View File
@@ -1815,25 +1815,73 @@ The last 100 journal lines for `ledmatrix.service` and
## Error tracking
Plugin errors are recorded by the display service (`ledmatrix.service`),
which runs the plugins. It publishes a snapshot to the shared cache directory
(`plugin_error_snapshot`) at most every 10 seconds, and only when something
changed, so these endpoints lag the display by up to about 15 seconds. The
counts cover the display service's current run: they start at zero when it
restarts. Error messages and stack traces have credentials redacted, and
messages, traces and context values are truncated in the snapshot.
Every response below adds three fields to the shape it always had:
| Field | Meaning |
|---|---|
| `snapshot_available` | `false` until the display service has reported (for example, it is not running). Counts are then zero. |
| `generated_at` | When the display service produced the snapshot (ISO, the Pi's local time), or `null`. |
| `clear_pending` | A clear has been requested and the display service has not applied it yet. |
### Get Error Summary
**GET** `/api/v3/errors/summary`
Aggregated counts, detected patterns and recent errors across plugins and
core components.
Aggregated counts, detected patterns and recent errors (the last 20).
```json
{
"status": "success",
"data": {
"session_start": "2026-09-23T09:40:02.118000",
"total_errors": 13,
"error_rate_per_hour": 41.2,
"error_counts_by_type": {"ConnectionError": 12, "ValueError": 1},
"plugin_error_counts": {"weather": {"ConnectionError": 12}, "stocks": {"ValueError": 1}},
"active_patterns": {
"ConnectionError": {
"error_type": "ConnectionError", "count": 12,
"first_seen": "2026-09-23T09:41:10.500000", "last_seen": "2026-09-23T09:58:36.020000",
"affected_plugins": ["weather"], "sample_messages": ["Read timed out."],
"severity": "error"
}
},
"recent_errors": [
{"error_type": "ValueError", "message": "could not parse price",
"timestamp": "2026-09-23T09:58:36.100000", "context": {},
"plugin_id": "stocks", "operation": "update", "stack_trace": "Traceback ..."}
],
"generated_at": "2026-09-23T09:58:40.000000",
"snapshot_available": true,
"clear_pending": false
},
"message": "Error summary retrieved"
}
```
### Get Plugin Errors
**GET** `/api/v3/errors/plugin/<plugin_id>`
Error health and statistics for one plugin.
Error health and statistics for one plugin: `plugin_id`, `status`
(`healthy`, `degraded` or `unhealthy`), `total_errors`, `error_types`,
`recent_error_count`, `last_error` (a `recent_errors` entry or `null`), plus
the three fields above. A plugin with no recorded errors is `healthy`.
### Clear Errors
**POST** `/api/v3/errors/clear`
Clear error records older than `max_age_hours` (default 24, 1-8760).
Returns `data.cleared_count`.
Clear error records older than `max_age_hours` (default 24, 1-8760), or every
error with `"all": true` (`max_age_hours` is then ignored).
```json
{
@@ -1841,6 +1889,32 @@ Returns `data.cleared_count`.
}
```
The clear is asynchronous. The web interface records a request
(`plugin_error_clear_request` in the shared cache), and the display service
applies it within about 5 seconds, rebuilding its counts from the errors it
keeps and republishing. Reads hide the cleared errors from the moment the
request is recorded. Until the display service applies an age-based clear,
`recent_errors` and `active_patterns` are already filtered but the counts
are the old ones, and `clear_pending` is `true`.
```json
{
"status": "success",
"data": {
"cleared_count": 13,
"clear_requested": true,
"request_id": "5f0c1e...",
"cutoff": "2026-09-23T09:59:02.310000"
},
"message": "Clear of all errors requested; the display service applies it within about 5 seconds"
}
```
`cleared_count` is how many of the reported errors the clear hides. It is
`null` when that cannot be known before the display service applies it (an
age-based clear over more errors than the report lists). A request that
could not be written to the shared cache answers `500`.
---
## Health and Status