fix(plugin-system): unload/update race, failed-load cleanup, limits validation, schema lookup, install rollback (#653)

* fix(plugin-system): unload/update race, failed-load module cleanup, limits validation, schema lookup, install rollback, op-queue dedupe

- unload_plugin takes the per-plugin lock (5s bounded) before cleanup(),
  and an update() that finishes after its plugin was unloaded no longer
  sets the state back to ENABLED.
- A load that fails after import drops plugin_<id> and its submodules
  and forgets its manager fonts, so a fixed plugin reloads new code.
- Resource limits are validated as non-negative numbers: 400 at
  POST /plugins/limits, bad cached records ignored with one warning.
  Route docstrings note health/metrics reset and limits only change the
  web process's view.
- SchemaManager.get_schema_path resolves each search dir via
  resolve_plugin_dir (manifest id, ledmatrix-<id>) before the literal
  paths; plugins/ still before plugin-repos/. Misses cached 30s and
  logged once at DEBUG.
- install_from_url sets an existing copy aside and restores it if the
  move fails, under the per-plugin reinstall lock.
- Operation queue refuses a second pending op for a plugin and trims
  _operations with history.
- get_vegas_render_width reads display_manager.width first.
- get_logger in store/schema/health/resource/saved_repositories;
  UTF-8 reads in store_manager and state_manager.
- Docs: update_interval precedence (manifest over config) stated where
  users are told to set it in config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): build the limits 400 message from the field name, not an exception

CodeQL flagged str(e) flowing into the response. invalid_limit_field()
returns the offending field without raising, and limits_from_dict uses it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-28 10:41:40 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent 0e9e2cabba
commit c00bf5e8e6
24 changed files with 869 additions and 49 deletions
+62 -3
View File
@@ -5,12 +5,14 @@ Tracks resource usage (memory, CPU, execution time) for plugins.
Provides resource limits and performance monitoring.
"""
import math
import time
import logging
import threading
from typing import Dict, Optional, Any, Callable
from dataclasses import dataclass, field, fields
from src.logging_config import get_logger
try:
import psutil
PSUTIL_AVAILABLE = True
@@ -31,6 +33,52 @@ class ResourceLimits:
warning_threshold: float = 0.8 # Warning at 80% of limit
_LIMIT_FIELDS = ('max_memory_mb', 'max_cpu_percent', 'max_execution_time',
'warning_threshold')
def invalid_limit_field(data: Any) -> Optional[str]:
"""The first field of a limits mapping that isn't a valid limit, or None.
``"limits"`` when ``data`` isn't a mapping at all. Separate from
limits_from_dict so a caller can report the problem without passing an
exception's text back to a client.
"""
if not isinstance(data, dict):
return 'limits'
for name in _LIMIT_FIELDS:
value = data.get(name)
if value is None:
continue
# bool is an int subclass; True is not a limit anyone meant.
if (isinstance(value, bool) or not isinstance(value, (int, float))
or not math.isfinite(value) or value < 0):
return name
return None
def limits_from_dict(data: Any) -> ResourceLimits:
"""Build ResourceLimits from a JSON-shaped mapping, validating each value.
A dataclass does not enforce its annotations, so ResourceLimits built from
raw request JSON or a cached record happily stores ``"50"`` -- and then
every monitored update() raises TypeError comparing a float with it. Each
``max_*`` value must be absent/None (no limit) or a non-negative number;
``warning_threshold`` defaults to 0.8. Unknown keys are ignored.
Raises:
ValueError: naming the first offending field.
"""
bad = invalid_limit_field(data)
if bad == 'limits':
raise ValueError(f"limits must be an object, got {type(data).__name__}")
if bad:
raise ValueError(
f"{bad} must be a non-negative number or null, got {data.get(bad)!r}")
return ResourceLimits(**{name: data[name] for name in _LIMIT_FIELDS
if data.get(name) is not None})
@dataclass
class ResourceMetrics:
"""Resource usage metrics for a plugin.
@@ -86,11 +134,12 @@ class PluginResourceMonitor:
"""
self.cache_manager = cache_manager
self.enable_monitoring = enable_monitoring and PSUTIL_AVAILABLE
self.logger = logging.getLogger(__name__)
self.logger = get_logger(__name__)
# Resource metrics per plugin
self._metrics: Dict[str, ResourceMetrics] = {}
self._limits: Dict[str, ResourceLimits] = {}
self._bad_limits_warned: set = set()
# When each plugin's metrics last reached the cache. Metrics change on
# every call, so they cannot be de-duplicated the way health state can;
# they are rate-limited instead. See _METRICS_PERSIST_INTERVAL.
@@ -230,7 +279,17 @@ class PluginResourceMonitor:
cache_key = self._get_limits_key(plugin_id)
cached = self.cache_manager.get(cache_key, max_age=None)
if cached:
self._limits[plugin_id] = ResourceLimits(**cached)
try:
self._limits[plugin_id] = limits_from_dict(cached)
except ValueError as e:
# Treat as no limits rather than letting every update
# of this plugin raise; warn once, not on every call.
if plugin_id not in self._bad_limits_warned:
self._bad_limits_warned.add(plugin_id)
self.logger.warning(
"Ignoring cached resource limits for %s: %s",
plugin_id, e)
return None
else:
return None
return self._limits[plugin_id]