mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 06:15:09 +00:00
fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator The error aggregator is a per-process singleton and only the display service runs plugins, so only its aggregator records anything. The routes read the web process's own, empty one and always reported no errors. The display service now publishes a bounded snapshot of its aggregator to the shared cache (plugin_error_snapshot) from a daemon thread: at most once every 10 s and only when something changed, never raising into the caller. The routes read it and keep their response shapes, adding snapshot_available, generated_at and clear_pending; exception text has credentials redacted. POST /errors/clear writes a clear request (plugin_error_clear_request) that the display applies on its next 5 s tick via the new clear_before(), which keeps errors recorded after the cutoff and rebuilds the counts. Until the snapshot acknowledges the request, reads hide everything before the cutoff, so a snapshot written just before the click cannot bring errors back. Adds "all": true; cleared_count is null when only the display can know it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): show plugin errors in the Logs tab A compact panel under the log viewer: per-plugin error counts, repeating errors (type, count, affected plugins, a sample message, last seen) and a Clear button, with empty states for "no errors" and "display service hasn't reported yet". Polls every 15 s while the tab is active; all text goes through escapeHtml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: describe where plugin error reports come from and how clear works Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(errors): redact the published snapshot before clipping it Keeping only a traceback's tail (or clipping a message) could cut an `api_key=` marker off while keeping the secret after it, and the web side's redaction would then have nothing to match. The display now redacts every free-text field of the snapshot first. The patterns move to a Flask-free src/redaction.py so the display service can use them; redact_text in the web error handler uses the same function, unchanged in behaviour. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1815,25 +1815,73 @@ The last 100 journal lines for `ledmatrix.service` and
|
||||
|
||||
## Error tracking
|
||||
|
||||
Plugin errors are recorded by the display service (`ledmatrix.service`),
|
||||
which runs the plugins. It publishes a snapshot to the shared cache directory
|
||||
(`plugin_error_snapshot`) at most every 10 seconds, and only when something
|
||||
changed, so these endpoints lag the display by up to about 15 seconds. The
|
||||
counts cover the display service's current run: they start at zero when it
|
||||
restarts. Error messages and stack traces have credentials redacted, and
|
||||
messages, traces and context values are truncated in the snapshot.
|
||||
|
||||
Every response below adds three fields to the shape it always had:
|
||||
|
||||
| Field | Meaning |
|
||||
|---|---|
|
||||
| `snapshot_available` | `false` until the display service has reported (for example, it is not running). Counts are then zero. |
|
||||
| `generated_at` | When the display service produced the snapshot (ISO, the Pi's local time), or `null`. |
|
||||
| `clear_pending` | A clear has been requested and the display service has not applied it yet. |
|
||||
|
||||
### Get Error Summary
|
||||
|
||||
**GET** `/api/v3/errors/summary`
|
||||
|
||||
Aggregated counts, detected patterns and recent errors across plugins and
|
||||
core components.
|
||||
Aggregated counts, detected patterns and recent errors (the last 20).
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "success",
|
||||
"data": {
|
||||
"session_start": "2026-09-23T09:40:02.118000",
|
||||
"total_errors": 13,
|
||||
"error_rate_per_hour": 41.2,
|
||||
"error_counts_by_type": {"ConnectionError": 12, "ValueError": 1},
|
||||
"plugin_error_counts": {"weather": {"ConnectionError": 12}, "stocks": {"ValueError": 1}},
|
||||
"active_patterns": {
|
||||
"ConnectionError": {
|
||||
"error_type": "ConnectionError", "count": 12,
|
||||
"first_seen": "2026-09-23T09:41:10.500000", "last_seen": "2026-09-23T09:58:36.020000",
|
||||
"affected_plugins": ["weather"], "sample_messages": ["Read timed out."],
|
||||
"severity": "error"
|
||||
}
|
||||
},
|
||||
"recent_errors": [
|
||||
{"error_type": "ValueError", "message": "could not parse price",
|
||||
"timestamp": "2026-09-23T09:58:36.100000", "context": {},
|
||||
"plugin_id": "stocks", "operation": "update", "stack_trace": "Traceback ..."}
|
||||
],
|
||||
"generated_at": "2026-09-23T09:58:40.000000",
|
||||
"snapshot_available": true,
|
||||
"clear_pending": false
|
||||
},
|
||||
"message": "Error summary retrieved"
|
||||
}
|
||||
```
|
||||
|
||||
### Get Plugin Errors
|
||||
|
||||
**GET** `/api/v3/errors/plugin/<plugin_id>`
|
||||
|
||||
Error health and statistics for one plugin.
|
||||
Error health and statistics for one plugin: `plugin_id`, `status`
|
||||
(`healthy`, `degraded` or `unhealthy`), `total_errors`, `error_types`,
|
||||
`recent_error_count`, `last_error` (a `recent_errors` entry or `null`), plus
|
||||
the three fields above. A plugin with no recorded errors is `healthy`.
|
||||
|
||||
### Clear Errors
|
||||
|
||||
**POST** `/api/v3/errors/clear`
|
||||
|
||||
Clear error records older than `max_age_hours` (default 24, 1-8760).
|
||||
Returns `data.cleared_count`.
|
||||
Clear error records older than `max_age_hours` (default 24, 1-8760), or every
|
||||
error with `"all": true` (`max_age_hours` is then ignored).
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -1841,6 +1889,32 @@ Returns `data.cleared_count`.
|
||||
}
|
||||
```
|
||||
|
||||
The clear is asynchronous. The web interface records a request
|
||||
(`plugin_error_clear_request` in the shared cache), and the display service
|
||||
applies it within about 5 seconds, rebuilding its counts from the errors it
|
||||
keeps and republishing. Reads hide the cleared errors from the moment the
|
||||
request is recorded. Until the display service applies an age-based clear,
|
||||
`recent_errors` and `active_patterns` are already filtered but the counts
|
||||
are the old ones, and `clear_pending` is `true`.
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "success",
|
||||
"data": {
|
||||
"cleared_count": 13,
|
||||
"clear_requested": true,
|
||||
"request_id": "5f0c1e...",
|
||||
"cutoff": "2026-09-23T09:59:02.310000"
|
||||
},
|
||||
"message": "Clear of all errors requested; the display service applies it within about 5 seconds"
|
||||
}
|
||||
```
|
||||
|
||||
`cleared_count` is how many of the reported errors the clear hides. It is
|
||||
`null` when that cannot be known before the display service applies it (an
|
||||
age-based clear over more errors than the report lists). A request that
|
||||
could not be written to the shared cache answers `500`.
|
||||
|
||||
---
|
||||
|
||||
## Health and Status
|
||||
|
||||
Reference in New Issue
Block a user