mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-08-01 08:48:05 +00:00
feat: activate dormant plugin health/metrics subsystem and surface it in the web UI (#388)
* feat(plugin-system): activate dormant plugin health & metrics subsystem
PluginManager shipped a fully-built health tracker, resource monitor and
circuit breaker that were never instantiated (health_tracker/resource_monitor
were left as None), so the circuit breaker never engaged and the existing
health/metrics API routes always returned "not available".
- DisplayController now wires a PluginHealthTracker and PluginResourceMonitor
onto the plugin manager, enabling the circuit breaker (a repeatedly-failing
plugin's update() is skipped after consecutive failures, then retried after
a cooldown) and per-plugin execution-time metrics. Both persist to the
shared cache.
- load_plugin() now validates each plugin's config against its JSON schema in
a strictly warn/degrade-only way: a violation logs a warning and flags the
plugin degraded in the health tracker, but never changes whether the plugin
loads or its pass/fail behaviour. Adds PluginHealthTracker.set_degraded(),
which never touches the circuit breaker.
- ResourceMonitor CPU/memory sampling now reuses a cached psutil.Process and
reads cpu_percent(interval=None), so monitoring no longer blocks ~100ms per
call on the display loop's update path.
- Fix DiskCache.get() raising TypeError for max_age=None ("never expires"),
which silently discarded persisted plugin health/metrics on read and thus
broke cross-process and post-restart surfacing.
- Fix two dead PluginManager helpers that called non-existent tracker methods.
Tests: new test_resource_monitor, test_plugin_health,
test_plugin_manager_schema_soft; extended test_cache_manager and
test_display_controller.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq
* feat(web-ui): surface plugin health, metrics and load state
With the health/metrics subsystem now active in the display service, expose it
in the web UI (which runs as a separate process from the display loop):
- Wire a health tracker / resource monitor backed by the shared on-disk cache
into the web process so /api/v3/plugins/health and /plugins/metrics read the
data the display service persists.
- Build those route responses per installed plugin id (the tracker's in-memory
view is empty in a fresh web process) so cross-process data is included.
- Add state + error_info to /plugins/installed entries so the UI can show why a
plugin isn't running instead of just loaded:false.
- Add a "Plugin Health" panel to the Tools page (circuit status, avg/max update
time, update count, last error) plus PluginAPI.getPluginMetrics().
Tests: route-level tests for the health/metrics endpoints in test_web_api.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq
* fix(plugin-metrics): refresh cross-process health/metrics reads; type hints
Addresses CodeRabbit review on #388:
- Major: the web process's health/resource trackers cached the first persisted
read in an in-memory dict (and the CacheManager memory tier held max_age=None
entries indefinitely), so a long-lived web process showed the first snapshot
and never reflected the display service's later updates. Add an opt-in
force_reload path (get_health_summary/get_health_state/_load_health_state and
get_metrics_summary/get_metrics) that bypasses the in-memory copy and, via a
new memory_ttl passthrough on CacheManager.get, the cache manager's memory
tier — so each /plugins/health and /plugins/metrics poll reads fresh persisted
state. Default behaviour (force_reload=False) is unchanged for the display
process and existing callers.
- Minor: DiskCache.get type hint is now Optional[int] with the None ("never
expires") semantics documented, matching MemoryCache.get.
Tests: new force_reload staleness cases in test_plugin_health and
test_resource_monitor.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq
---------
Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
Vendored
+12
-5
@@ -68,14 +68,15 @@ class DiskCache:
|
||||
return None
|
||||
return os.path.join(self.cache_dir, f"{key}.json")
|
||||
|
||||
def get(self, key: str, max_age: int = 300) -> Optional[Dict[str, Any]]:
|
||||
def get(self, key: str, max_age: Optional[int] = 300) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Get data from disk cache.
|
||||
|
||||
|
||||
Args:
|
||||
key: Cache key
|
||||
max_age: Maximum age in seconds
|
||||
|
||||
max_age: Maximum age in seconds; None disables age-based expiry
|
||||
(the record never counts as stale). Mirrors MemoryCache.get.
|
||||
|
||||
Returns:
|
||||
Cached data or None if not found or expired
|
||||
"""
|
||||
@@ -105,7 +106,13 @@ class DiskCache:
|
||||
record_ts = None
|
||||
|
||||
now = time.time()
|
||||
if record_ts is None or (now - record_ts) <= max_age:
|
||||
# max_age=None means "never expires" (mirrors MemoryCache and the
|
||||
# cache_manager docstring). Guard it explicitly — otherwise the
|
||||
# comparison below raises TypeError and the record is treated as a
|
||||
# miss, which silently breaks callers that persist long-lived state
|
||||
# via get(key, max_age=None) (e.g. plugin health/metrics that must
|
||||
# survive restarts and be read cross-process).
|
||||
if record_ts is None or max_age is None or (now - record_ts) <= max_age:
|
||||
return record
|
||||
else:
|
||||
# Stale on disk; keep file for potential diagnostics but treat as miss
|
||||
|
||||
+13
-3
@@ -574,9 +574,19 @@ class CacheManager:
|
||||
}
|
||||
return self.save_cache(data_type, cache_data)
|
||||
|
||||
def get(self, key: str, max_age: int = 300) -> Optional[Dict[str, Any]]:
|
||||
"""Get data from cache if it exists and is not stale."""
|
||||
cached_data = self.get_cached_data(key, max_age)
|
||||
def get(self, key: str, max_age: Optional[int] = 300,
|
||||
memory_ttl: Optional[int] = None) -> Optional[Dict[str, Any]]:
|
||||
"""Get data from cache if it exists and is not stale.
|
||||
|
||||
Args:
|
||||
key: Cache key
|
||||
max_age: Max age (seconds) for the on-disk entry; None never expires.
|
||||
memory_ttl: Max age (seconds) for the in-memory entry. Pass 0 to
|
||||
bypass the memory tier and force a fresh read from disk — used by
|
||||
cross-process readers that must observe another process's latest
|
||||
write rather than a stale first snapshot. Defaults to max_age.
|
||||
"""
|
||||
cached_data = self.get_cached_data(key, max_age, memory_ttl=memory_ttl)
|
||||
if cached_data and 'data' in cached_data:
|
||||
return cached_data['data']
|
||||
return cached_data
|
||||
|
||||
@@ -230,7 +230,24 @@ class DisplayController:
|
||||
cache_manager=self.cache_manager,
|
||||
font_manager=self.font_manager
|
||||
)
|
||||
|
||||
|
||||
# Activate the plugin health/metrics subsystem. PluginManager leaves
|
||||
# health_tracker/resource_monitor as None by default; wiring real
|
||||
# instances here turns on the circuit breaker (a repeatedly-failing
|
||||
# plugin's update() is skipped after consecutive failures, then
|
||||
# retried after a cooldown) and per-plugin execution-time metrics.
|
||||
# Both persist to the shared cache so the web UI can surface them.
|
||||
# Done before discovery/loading so load-time schema warnings have a
|
||||
# tracker to record against.
|
||||
try:
|
||||
from src.plugin_system.plugin_health import PluginHealthTracker
|
||||
from src.plugin_system.resource_monitor import PluginResourceMonitor
|
||||
self.plugin_manager.health_tracker = PluginHealthTracker(self.cache_manager)
|
||||
self.plugin_manager.resource_monitor = PluginResourceMonitor(self.cache_manager)
|
||||
logger.info("Plugin health tracking and resource monitoring enabled")
|
||||
except Exception as e:
|
||||
logger.warning("Could not enable plugin health/resource monitoring: %s", e)
|
||||
|
||||
# Validate plugins after plugin manager is created
|
||||
try:
|
||||
from src.startup_validator import StartupValidator
|
||||
|
||||
@@ -52,11 +52,18 @@ class PluginHealthTracker:
|
||||
"""Get cache key for plugin health data."""
|
||||
return f"plugin_health:{plugin_id}"
|
||||
|
||||
def _load_health_state(self, plugin_id: str) -> Dict[str, Any]:
|
||||
"""Load health state from cache or return defaults."""
|
||||
def _load_health_state(self, plugin_id: str, force_reload: bool = False) -> Dict[str, Any]:
|
||||
"""Load health state from cache or return defaults.
|
||||
|
||||
``force_reload=True`` bypasses the cache manager's in-memory tier so a
|
||||
read-only consumer (e.g. the web process) observes the writer process's
|
||||
latest persisted state instead of a stale first snapshot.
|
||||
"""
|
||||
cache_key = self._get_health_key(plugin_id)
|
||||
cached = self.cache_manager.get(cache_key, max_age=None)
|
||||
|
||||
cached = self.cache_manager.get(
|
||||
cache_key, max_age=None, memory_ttl=0 if force_reload else None
|
||||
)
|
||||
|
||||
if cached:
|
||||
return cached
|
||||
|
||||
@@ -79,10 +86,17 @@ class PluginHealthTracker:
|
||||
self.cache_manager.set(cache_key, state) # Persist indefinitely
|
||||
self._health_state[plugin_id] = state
|
||||
|
||||
def get_health_state(self, plugin_id: str) -> Dict[str, Any]:
|
||||
"""Get current health state for a plugin."""
|
||||
if plugin_id not in self._health_state:
|
||||
self._health_state[plugin_id] = self._load_health_state(plugin_id)
|
||||
def get_health_state(self, plugin_id: str, force_reload: bool = False) -> Dict[str, Any]:
|
||||
"""Get current health state for a plugin.
|
||||
|
||||
``force_reload=True`` re-reads the persisted state from the cache,
|
||||
bypassing the in-memory copy — needed by cross-process readers that
|
||||
would otherwise be pinned to the first snapshot they loaded.
|
||||
"""
|
||||
if force_reload or plugin_id not in self._health_state:
|
||||
self._health_state[plugin_id] = self._load_health_state(
|
||||
plugin_id, force_reload=force_reload
|
||||
)
|
||||
return self._health_state[plugin_id]
|
||||
|
||||
def record_success(self, plugin_id: str) -> None:
|
||||
@@ -139,6 +153,28 @@ class PluginHealthTracker:
|
||||
|
||||
self._save_health_state(plugin_id, state)
|
||||
|
||||
def set_degraded(self, plugin_id: str, reason: Optional[str]) -> None:
|
||||
"""Flag (or clear) a plugin as degraded without touching the circuit breaker.
|
||||
|
||||
Used for non-fatal issues — e.g. a config that no longer satisfies the
|
||||
plugin's schema — that should be surfaced to the user but must NOT cause
|
||||
the plugin to be skipped or counted as a runtime failure. Passing
|
||||
``reason=None`` clears the flag. The write is skipped when nothing
|
||||
actually changes, so calling this on every load is cheap.
|
||||
|
||||
Args:
|
||||
plugin_id: Plugin identifier
|
||||
reason: Human-readable reason string, or None to clear the flag
|
||||
"""
|
||||
state = self.get_health_state(plugin_id)
|
||||
new_degraded = bool(reason)
|
||||
new_reason = reason if reason else None
|
||||
if state.get('degraded', False) == new_degraded and state.get('degraded_reason') == new_reason:
|
||||
return # No change — avoid a redundant cache write
|
||||
state['degraded'] = new_degraded
|
||||
state['degraded_reason'] = new_reason
|
||||
self._save_health_state(plugin_id, state)
|
||||
|
||||
def should_skip_plugin(self, plugin_id: str) -> bool:
|
||||
"""
|
||||
Check if plugin should be skipped due to circuit breaker.
|
||||
@@ -181,9 +217,13 @@ class PluginHealthTracker:
|
||||
|
||||
return False
|
||||
|
||||
def get_health_summary(self, plugin_id: str) -> Dict[str, Any]:
|
||||
"""Get health summary for a plugin."""
|
||||
state = self.get_health_state(plugin_id)
|
||||
def get_health_summary(self, plugin_id: str, force_reload: bool = False) -> Dict[str, Any]:
|
||||
"""Get health summary for a plugin.
|
||||
|
||||
``force_reload=True`` refreshes from the persisted cache first so
|
||||
cross-process readers reflect the writer's latest state.
|
||||
"""
|
||||
state = self.get_health_state(plugin_id, force_reload=force_reload)
|
||||
|
||||
total_calls = state.get('total_successes', 0) + state.get('total_failures', 0)
|
||||
success_rate = 0.0
|
||||
@@ -201,6 +241,8 @@ class PluginHealthTracker:
|
||||
'last_failure_time': state.get('last_failure_time'),
|
||||
'last_error': state.get('last_error'),
|
||||
'is_healthy': state.get('circuit_state') == CircuitState.CLOSED.value,
|
||||
'degraded': state.get('degraded', False),
|
||||
'degraded_reason': state.get('degraded_reason'),
|
||||
'circuit_opened_time': state.get('circuit_opened_time'),
|
||||
'half_open_start_time': state.get('half_open_start_time')
|
||||
}
|
||||
|
||||
@@ -390,7 +390,15 @@ class PluginManager:
|
||||
self.logger.error("Error validating plugin %s config: %s", plugin_id, e, exc_info=True)
|
||||
self.state_manager.set_state(plugin_id, PluginState.ERROR, error=e)
|
||||
return False
|
||||
|
||||
|
||||
# Schema validation (warn/degrade only — never blocks loading).
|
||||
# A config that violates the plugin's JSON schema is surfaced to the
|
||||
# user (log warning + degraded flag in the health tracker) but the
|
||||
# plugin still loads exactly as it does today. This deliberately does
|
||||
# NOT change load_plugin()'s pass/fail behaviour for any plugin that
|
||||
# loads under the current code.
|
||||
self._validate_config_schema_soft(plugin_id, config)
|
||||
|
||||
# Store plugin instance
|
||||
self.plugins[plugin_id] = plugin_instance
|
||||
self.plugin_last_update[plugin_id] = 0.0
|
||||
@@ -419,6 +427,59 @@ class PluginManager:
|
||||
self.state_manager.set_state(plugin_id, PluginState.ERROR, error=e)
|
||||
return False
|
||||
|
||||
def _validate_config_schema_soft(self, plugin_id: str, config: Dict[str, Any]) -> None:
|
||||
"""Validate a plugin's config against its JSON schema — warn/degrade only.
|
||||
|
||||
On a schema violation this logs a warning and marks the plugin degraded
|
||||
in the health tracker (when one is wired), so the problem is visible in
|
||||
the web UI. It never raises, never changes plugin state, and never
|
||||
affects whether the plugin loads. ``config`` here has already been
|
||||
merged with schema defaults by the caller, so fields that ship a default
|
||||
never appear "missing" — only genuinely user-supplied required fields
|
||||
(e.g. an API key) can trip the required-field check.
|
||||
"""
|
||||
try:
|
||||
schema = self.schema_manager.load_schema(plugin_id)
|
||||
except Exception as e: # pragma: no cover - defensive
|
||||
self.logger.debug("Could not load schema for %s: %s", plugin_id, e)
|
||||
return
|
||||
|
||||
if not schema:
|
||||
# No schema shipped — nothing to validate. Clear any stale flag.
|
||||
self._set_degraded_safe(plugin_id, None)
|
||||
return
|
||||
|
||||
try:
|
||||
is_valid, errors = self.schema_manager.validate_config_against_schema(
|
||||
config, schema, plugin_id
|
||||
)
|
||||
except Exception as e: # pragma: no cover - defensive
|
||||
# Validation machinery itself failed — do not penalise the plugin.
|
||||
self.logger.debug("Schema validation raised for %s: %s", plugin_id, e)
|
||||
return
|
||||
|
||||
if is_valid or not errors:
|
||||
self._set_degraded_safe(plugin_id, None)
|
||||
return
|
||||
|
||||
summary = "; ".join(errors[:5])
|
||||
if len(errors) > 5:
|
||||
summary += f" (+{len(errors) - 5} more)"
|
||||
self.logger.warning(
|
||||
"Plugin %s config does not match its schema (loading anyway): %s",
|
||||
plugin_id, summary,
|
||||
)
|
||||
self._set_degraded_safe(plugin_id, f"Config schema: {summary}")
|
||||
|
||||
def _set_degraded_safe(self, plugin_id: str, reason: Optional[str]) -> None:
|
||||
"""Best-effort ``health_tracker.set_degraded`` that never raises."""
|
||||
if not self.health_tracker:
|
||||
return
|
||||
try:
|
||||
self.health_tracker.set_degraded(plugin_id, reason)
|
||||
except Exception as e: # pragma: no cover - defensive
|
||||
self.logger.debug("Could not set degraded flag for %s: %s", plugin_id, e)
|
||||
|
||||
def unload_plugin(self, plugin_id: str) -> bool:
|
||||
"""
|
||||
Unload a plugin by ID.
|
||||
@@ -836,7 +897,7 @@ class PluginManager:
|
||||
|
||||
# Get health tracker metrics if available
|
||||
if self.health_tracker:
|
||||
health_info = self.health_tracker.get_plugin_health(plugin_id)
|
||||
health_info = self.health_tracker.get_health_summary(plugin_id)
|
||||
plugin_metrics['health'] = health_info
|
||||
else:
|
||||
plugin_metrics['health'] = {'status': 'unknown'}
|
||||
@@ -861,7 +922,7 @@ class PluginManager:
|
||||
|
||||
# Get resource monitor metrics if available
|
||||
if self.resource_monitor:
|
||||
resource_info = self.resource_monitor.get_plugin_metrics(plugin_id)
|
||||
resource_info = self.resource_monitor.get_metrics_summary(plugin_id)
|
||||
plugin_metrics['resources'] = resource_info
|
||||
else:
|
||||
plugin_metrics['resources'] = {'status': 'unknown'}
|
||||
|
||||
@@ -71,17 +71,32 @@ class PluginResourceMonitor:
|
||||
self.cache_manager = cache_manager
|
||||
self.enable_monitoring = enable_monitoring and PSUTIL_AVAILABLE
|
||||
self.logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
# Resource metrics per plugin
|
||||
self._metrics: Dict[str, ResourceMetrics] = {}
|
||||
self._limits: Dict[str, ResourceLimits] = {}
|
||||
|
||||
|
||||
# Thread-local storage for execution tracking
|
||||
self._local = threading.local()
|
||||
|
||||
|
||||
# Lock for thread-safe access
|
||||
self._lock = threading.Lock()
|
||||
|
||||
|
||||
# Cache a single psutil.Process handle. Reusing the same handle is what
|
||||
# lets cpu_percent() be read non-blocking (interval=None): psutil returns
|
||||
# the utilisation since the *previous* call on that same object. Creating
|
||||
# a fresh Process() per call would force interval-based sampling that
|
||||
# blocks the caller — unacceptable on the display loop's update path.
|
||||
self._process = None
|
||||
if self.enable_monitoring:
|
||||
try:
|
||||
self._process = psutil.Process()
|
||||
# Prime cpu_percent so the first real measurement returns a
|
||||
# meaningful delta instead of 0.0.
|
||||
self._process.cpu_percent(interval=None)
|
||||
except Exception: # pragma: no cover - psutil edge cases
|
||||
self._process = None
|
||||
|
||||
if not PSUTIL_AVAILABLE and enable_monitoring:
|
||||
self.logger.warning(
|
||||
"psutil not available - resource monitoring will be limited to execution time only"
|
||||
@@ -95,13 +110,21 @@ class PluginResourceMonitor:
|
||||
"""Get cache key for plugin limits."""
|
||||
return f"plugin_limits:{plugin_id}"
|
||||
|
||||
def get_metrics(self, plugin_id: str) -> ResourceMetrics:
|
||||
"""Get current metrics for a plugin."""
|
||||
def get_metrics(self, plugin_id: str, force_reload: bool = False) -> ResourceMetrics:
|
||||
"""Get current metrics for a plugin.
|
||||
|
||||
``force_reload=True`` bypasses both the in-memory copy and the cache
|
||||
manager's memory tier so a read-only consumer (e.g. the web process)
|
||||
sees the writer process's latest persisted metrics rather than a stale
|
||||
first snapshot.
|
||||
"""
|
||||
with self._lock:
|
||||
if plugin_id not in self._metrics:
|
||||
if force_reload or plugin_id not in self._metrics:
|
||||
# Try to load from cache
|
||||
cache_key = self._get_metrics_key(plugin_id)
|
||||
cached = self.cache_manager.get(cache_key, max_age=None)
|
||||
cached = self.cache_manager.get(
|
||||
cache_key, max_age=None, memory_ttl=0 if force_reload else None
|
||||
)
|
||||
if cached:
|
||||
metrics = ResourceMetrics(**cached)
|
||||
else:
|
||||
@@ -137,21 +160,24 @@ class PluginResourceMonitor:
|
||||
|
||||
def _get_process_memory_mb(self) -> float:
|
||||
"""Get current process memory usage in MB."""
|
||||
if not self.enable_monitoring:
|
||||
if not self.enable_monitoring or self._process is None:
|
||||
return 0.0
|
||||
try:
|
||||
process = psutil.Process()
|
||||
return process.memory_info().rss / 1024 / 1024
|
||||
return self._process.memory_info().rss / 1024 / 1024
|
||||
except Exception:
|
||||
return 0.0
|
||||
|
||||
def _get_process_cpu_percent(self, interval: float = 0.1) -> float:
|
||||
"""Get current process CPU usage percentage."""
|
||||
if not self.enable_monitoring:
|
||||
|
||||
def _get_process_cpu_percent(self) -> float:
|
||||
"""Get current process CPU usage percentage (non-blocking).
|
||||
|
||||
Reads cpu_percent(interval=None) against the cached process handle, so
|
||||
it returns immediately with the utilisation observed since the previous
|
||||
call rather than blocking to sample a fresh interval.
|
||||
"""
|
||||
if not self.enable_monitoring or self._process is None:
|
||||
return 0.0
|
||||
try:
|
||||
process = psutil.Process()
|
||||
return process.cpu_percent(interval=interval)
|
||||
return self._process.cpu_percent(interval=None)
|
||||
except Exception:
|
||||
return 0.0
|
||||
|
||||
@@ -281,9 +307,13 @@ class PluginResourceMonitor:
|
||||
self.logger.error(error_msg)
|
||||
raise ResourceLimitExceeded(error_msg)
|
||||
|
||||
def get_metrics_summary(self, plugin_id: str) -> Dict[str, Any]:
|
||||
"""Get metrics summary for a plugin."""
|
||||
metrics = self.get_metrics(plugin_id)
|
||||
def get_metrics_summary(self, plugin_id: str, force_reload: bool = False) -> Dict[str, Any]:
|
||||
"""Get metrics summary for a plugin.
|
||||
|
||||
``force_reload=True`` refreshes from the persisted cache first so
|
||||
cross-process readers reflect the writer's latest metrics.
|
||||
"""
|
||||
metrics = self.get_metrics(plugin_id, force_reload=force_reload)
|
||||
limits = self.get_limits(plugin_id)
|
||||
|
||||
avg_execution_time = 0.0
|
||||
|
||||
Reference in New Issue
Block a user