Files
LEDMatrix/docs/PLUGIN_ERROR_HANDLING.md
T
ChuckandClaude Opus 5.5 4a1fd7464a fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator

The error aggregator is a per-process singleton and only the display
service runs plugins, so only its aggregator records anything. The routes
read the web process's own, empty one and always reported no errors.

The display service now publishes a bounded snapshot of its aggregator to
the shared cache (plugin_error_snapshot) from a daemon thread: at most once
every 10 s and only when something changed, never raising into the caller.
The routes read it and keep their response shapes, adding
snapshot_available, generated_at and clear_pending; exception text has
credentials redacted.

POST /errors/clear writes a clear request (plugin_error_clear_request) that
the display applies on its next 5 s tick via the new clear_before(), which
keeps errors recorded after the cutoff and rebuilds the counts. Until the
snapshot acknowledges the request, reads hide everything before the cutoff,
so a snapshot written just before the click cannot bring errors back. Adds
"all": true; cleared_count is null when only the display can know it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): show plugin errors in the Logs tab

A compact panel under the log viewer: per-plugin error counts, repeating
errors (type, count, affected plugins, a sample message, last seen) and a
Clear button, with empty states for "no errors" and "display service
hasn't reported yet". Polls every 15 s while the tab is active; all text
goes through escapeHtml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe where plugin error reports come from and how clear works

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(errors): redact the published snapshot before clipping it

Keeping only a traceback's tail (or clipping a message) could cut an
`api_key=` marker off while keeping the secret after it, and the web side's
redaction would then have nothing to match. The display now redacts every
free-text field of the snapshot first. The patterns move to a Flask-free
src/redaction.py so the display service can use them; redact_text in the web
error handler uses the same function, unchanged in behaviour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:32:02 -04:00

258 lines
7.3 KiB
Markdown

# Plugin Error Handling Guide
This guide covers best practices for error handling in LEDMatrix plugins.
## Custom Exception Hierarchy
LEDMatrix provides typed exceptions for different error categories. Use these instead of generic `Exception`:
```python
from src.exceptions import PluginError, ConfigError, CacheError, DisplayError
# Plugin-related errors
raise PluginError("Failed to fetch data", plugin_id=self.plugin_id, context={"api": "ESPN"})
# Configuration errors
raise ConfigError("Invalid API key format", field="api_key")
# Cache errors
raise CacheError("Cache write failed", cache_key="game_data")
# Display errors
raise DisplayError("Failed to render", display_mode="live")
```
### Exception Context
All LEDMatrix exceptions support a `context` dict for additional debugging info:
```python
raise PluginError(
"API request failed",
plugin_id=self.plugin_id,
context={
"url": api_url,
"status_code": response.status_code,
"retry_count": 3
}
)
```
## Logging Best Practices
### Use the Plugin Logger
Every plugin has access to `self.logger`:
```python
class MyPlugin(BasePlugin):
def update(self):
self.logger.info("Starting data fetch")
self.logger.debug("API URL: %s", api_url)
self.logger.warning("Rate limit approaching")
self.logger.error("API request failed", exc_info=True)
```
### Log Levels
- **DEBUG**: Detailed info for troubleshooting (API URLs, parsed data)
- **INFO**: Normal operation milestones (plugin loaded, data fetched)
- **WARNING**: Recoverable issues (rate limits, cache miss, fallback used)
- **ERROR**: Failures that need attention (API down, display error)
### Include exc_info for Exceptions
```python
try:
response = requests.get(url)
except requests.RequestException as e:
self.logger.error("API request failed: %s", e, exc_info=True)
```
## Error Handling Patterns
### Never Use Bare except
```python
# BAD - swallows all errors including KeyboardInterrupt
try:
self.fetch_data()
except:
pass
# GOOD - catch specific exceptions
try:
self.fetch_data()
except requests.RequestException as e:
self.logger.warning("Network error, using cached data: %s", e)
self.data = self.get_cached_data()
```
### Graceful Degradation
```python
def update(self):
try:
self.data = self.fetch_live_data()
except requests.RequestException as e:
self.logger.warning("Live data unavailable: %s", e)
# Fall back to cache
cached = self.cache_manager.get(self.cache_key)
if cached:
self.logger.info("Using cached data")
self.data = cached
else:
self.logger.error("No cached data available")
self.data = None
```
### Validate Configuration Early
```python
def validate_config(self) -> bool:
"""Validate configuration at load time."""
api_key = self.config.get("api_key")
if not api_key:
self.logger.error("api_key is required but not configured")
return False
if not isinstance(api_key, str) or len(api_key) < 10:
self.logger.error("api_key appears to be invalid")
return False
return True
```
### Handle Display Errors
```python
def display(self, force_clear: bool = False) -> bool:
if not self.data:
if force_clear:
self.display_manager.clear()
self.display_manager.update_display()
return False
try:
self._render_content()
return True
except Exception as e:
self.logger.error("Display error: %s", e, exc_info=True)
# Clear display on error to prevent stale content
self.display_manager.clear()
self.display_manager.update_display()
return False
```
## Error Aggregation
LEDMatrix automatically tracks plugin errors: every exception or timeout from
a plugin's `update()` or `display()` is recorded by the display service,
which runs the plugins. See them in the web interface under **Logs → Plugin
errors**, or through the API:
```bash
# Get error summary
curl http://localhost:5000/api/v3/errors/summary
# Get plugin-specific health
curl http://localhost:5000/api/v3/errors/plugin/my-plugin
# Clear errors older than 24 hours (the default), or all of them
curl -X POST http://localhost:5000/api/v3/errors/clear
curl -X POST -H 'Content-Type: application/json' -d '{"all": true}' \
http://localhost:5000/api/v3/errors/clear
```
The web interface is a separate process, so it reads a snapshot the display
service writes to the shared cache directory (`plugin_error_snapshot`): at most
every 10 seconds, and only when something changed. Expect the numbers to lag
by up to about 15 seconds, and to start from zero when the display service
restarts. `snapshot_available` is `false` until the display service has
reported. A clear is a request the display service applies within about 5
seconds; the API hides the cleared errors immediately. Details and response
shapes: [REST API reference](REST_API_REFERENCE.md#error-tracking).
### Error Patterns
When the same error occurs repeatedly (5+ times in 60 minutes), it's detected as a pattern and logged as a warning. This helps identify systemic issues.
## Common Error Scenarios
### API Rate Limiting
```python
def fetch_data(self):
try:
response = requests.get(self.api_url)
if response.status_code == 429:
retry_after = int(response.headers.get("Retry-After", 60))
self.logger.warning("Rate limited, retry after %ds", retry_after)
self._rate_limited_until = time.time() + retry_after
return None
response.raise_for_status()
return response.json()
except requests.RequestException as e:
self.logger.error("API error: %s", e)
return None
```
### Timeout Handling
```python
def fetch_data(self):
try:
response = requests.get(self.api_url, timeout=10)
return response.json()
except requests.Timeout:
self.logger.warning("Request timed out, will retry next update")
return None
except requests.RequestException as e:
self.logger.error("Request failed: %s", e)
return None
```
### Missing Data Gracefully
```python
def get_team_logo(self, team_id):
logo_path = self.logos_dir / f"{team_id}.png"
if not logo_path.exists():
self.logger.debug("Logo not found for team %s, using default", team_id)
return self.default_logo
return Image.open(logo_path)
```
## Testing Error Handling
```python
def test_handles_api_error(mock_requests):
"""Test plugin handles API errors gracefully."""
mock_requests.get.side_effect = requests.RequestException("Network error")
plugin = MyPlugin(...)
plugin.update()
# Should not raise, should log warning, should have no data
assert plugin.data is None
def test_handles_invalid_json(mock_requests):
"""Test plugin handles invalid JSON response."""
mock_requests.get.return_value.json.side_effect = ValueError("Invalid JSON")
plugin = MyPlugin(...)
plugin.update()
assert plugin.data is None
```
## Checklist
- [ ] No bare `except:` clauses
- [ ] All exceptions logged with appropriate level
- [ ] `exc_info=True` for error-level logs
- [ ] Graceful degradation with cache fallbacks
- [ ] Configuration validated in `validate_config()`
- [ ] Display clears on error to prevent stale content
- [ ] Timeouts configured for all network requests