feat: activate dormant plugin health/metrics subsystem and surface it in the web UI (#388)

* feat(plugin-system): activate dormant plugin health & metrics subsystem

PluginManager shipped a fully-built health tracker, resource monitor and
circuit breaker that were never instantiated (health_tracker/resource_monitor
were left as None), so the circuit breaker never engaged and the existing
health/metrics API routes always returned "not available".

- DisplayController now wires a PluginHealthTracker and PluginResourceMonitor
  onto the plugin manager, enabling the circuit breaker (a repeatedly-failing
  plugin's update() is skipped after consecutive failures, then retried after
  a cooldown) and per-plugin execution-time metrics. Both persist to the
  shared cache.
- load_plugin() now validates each plugin's config against its JSON schema in
  a strictly warn/degrade-only way: a violation logs a warning and flags the
  plugin degraded in the health tracker, but never changes whether the plugin
  loads or its pass/fail behaviour. Adds PluginHealthTracker.set_degraded(),
  which never touches the circuit breaker.
- ResourceMonitor CPU/memory sampling now reuses a cached psutil.Process and
  reads cpu_percent(interval=None), so monitoring no longer blocks ~100ms per
  call on the display loop's update path.
- Fix DiskCache.get() raising TypeError for max_age=None ("never expires"),
  which silently discarded persisted plugin health/metrics on read and thus
  broke cross-process and post-restart surfacing.
- Fix two dead PluginManager helpers that called non-existent tracker methods.

Tests: new test_resource_monitor, test_plugin_health,
test_plugin_manager_schema_soft; extended test_cache_manager and
test_display_controller.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

* feat(web-ui): surface plugin health, metrics and load state

With the health/metrics subsystem now active in the display service, expose it
in the web UI (which runs as a separate process from the display loop):

- Wire a health tracker / resource monitor backed by the shared on-disk cache
  into the web process so /api/v3/plugins/health and /plugins/metrics read the
  data the display service persists.
- Build those route responses per installed plugin id (the tracker's in-memory
  view is empty in a fresh web process) so cross-process data is included.
- Add state + error_info to /plugins/installed entries so the UI can show why a
  plugin isn't running instead of just loaded:false.
- Add a "Plugin Health" panel to the Tools page (circuit status, avg/max update
  time, update count, last error) plus PluginAPI.getPluginMetrics().

Tests: route-level tests for the health/metrics endpoints in test_web_api.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

* fix(plugin-metrics): refresh cross-process health/metrics reads; type hints

Addresses CodeRabbit review on #388:

- Major: the web process's health/resource trackers cached the first persisted
  read in an in-memory dict (and the CacheManager memory tier held max_age=None
  entries indefinitely), so a long-lived web process showed the first snapshot
  and never reflected the display service's later updates. Add an opt-in
  force_reload path (get_health_summary/get_health_state/_load_health_state and
  get_metrics_summary/get_metrics) that bypasses the in-memory copy and, via a
  new memory_ttl passthrough on CacheManager.get, the cache manager's memory
  tier — so each /plugins/health and /plugins/metrics poll reads fresh persisted
  state. Default behaviour (force_reload=False) is unchanged for the display
  process and existing callers.
- Minor: DiskCache.get type hint is now Optional[int] with the None ("never
  expires") semantics documented, matching MemoryCache.get.

Tests: new force_reload staleness cases in test_plugin_health and
test_resource_monitor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

---------

Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-07-09 07:54:18 -04:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 85d321cf33
commit 3b93024993
16 changed files with 797 additions and 49 deletions
+64 -3
View File
@@ -390,7 +390,15 @@ class PluginManager:
self.logger.error("Error validating plugin %s config: %s", plugin_id, e, exc_info=True)
self.state_manager.set_state(plugin_id, PluginState.ERROR, error=e)
return False
# Schema validation (warn/degrade only — never blocks loading).
# A config that violates the plugin's JSON schema is surfaced to the
# user (log warning + degraded flag in the health tracker) but the
# plugin still loads exactly as it does today. This deliberately does
# NOT change load_plugin()'s pass/fail behaviour for any plugin that
# loads under the current code.
self._validate_config_schema_soft(plugin_id, config)
# Store plugin instance
self.plugins[plugin_id] = plugin_instance
self.plugin_last_update[plugin_id] = 0.0
@@ -419,6 +427,59 @@ class PluginManager:
self.state_manager.set_state(plugin_id, PluginState.ERROR, error=e)
return False
def _validate_config_schema_soft(self, plugin_id: str, config: Dict[str, Any]) -> None:
"""Validate a plugin's config against its JSON schema — warn/degrade only.
On a schema violation this logs a warning and marks the plugin degraded
in the health tracker (when one is wired), so the problem is visible in
the web UI. It never raises, never changes plugin state, and never
affects whether the plugin loads. ``config`` here has already been
merged with schema defaults by the caller, so fields that ship a default
never appear "missing" — only genuinely user-supplied required fields
(e.g. an API key) can trip the required-field check.
"""
try:
schema = self.schema_manager.load_schema(plugin_id)
except Exception as e: # pragma: no cover - defensive
self.logger.debug("Could not load schema for %s: %s", plugin_id, e)
return
if not schema:
# No schema shipped — nothing to validate. Clear any stale flag.
self._set_degraded_safe(plugin_id, None)
return
try:
is_valid, errors = self.schema_manager.validate_config_against_schema(
config, schema, plugin_id
)
except Exception as e: # pragma: no cover - defensive
# Validation machinery itself failed — do not penalise the plugin.
self.logger.debug("Schema validation raised for %s: %s", plugin_id, e)
return
if is_valid or not errors:
self._set_degraded_safe(plugin_id, None)
return
summary = "; ".join(errors[:5])
if len(errors) > 5:
summary += f" (+{len(errors) - 5} more)"
self.logger.warning(
"Plugin %s config does not match its schema (loading anyway): %s",
plugin_id, summary,
)
self._set_degraded_safe(plugin_id, f"Config schema: {summary}")
def _set_degraded_safe(self, plugin_id: str, reason: Optional[str]) -> None:
"""Best-effort ``health_tracker.set_degraded`` that never raises."""
if not self.health_tracker:
return
try:
self.health_tracker.set_degraded(plugin_id, reason)
except Exception as e: # pragma: no cover - defensive
self.logger.debug("Could not set degraded flag for %s: %s", plugin_id, e)
def unload_plugin(self, plugin_id: str) -> bool:
"""
Unload a plugin by ID.
@@ -836,7 +897,7 @@ class PluginManager:
# Get health tracker metrics if available
if self.health_tracker:
health_info = self.health_tracker.get_plugin_health(plugin_id)
health_info = self.health_tracker.get_health_summary(plugin_id)
plugin_metrics['health'] = health_info
else:
plugin_metrics['health'] = {'status': 'unknown'}
@@ -861,7 +922,7 @@ class PluginManager:
# Get resource monitor metrics if available
if self.resource_monitor:
resource_info = self.resource_monitor.get_plugin_metrics(plugin_id)
resource_info = self.resource_monitor.get_metrics_summary(plugin_id)
plugin_metrics['resources'] = resource_info
else:
plugin_metrics['resources'] = {'status': 'unknown'}