fix(plugin-system): unload/update race, failed-load cleanup, limits validation, schema lookup, install rollback (#653)

* fix(plugin-system): unload/update race, failed-load module cleanup, limits validation, schema lookup, install rollback, op-queue dedupe

- unload_plugin takes the per-plugin lock (5s bounded) before cleanup(),
  and an update() that finishes after its plugin was unloaded no longer
  sets the state back to ENABLED.
- A load that fails after import drops plugin_<id> and its submodules
  and forgets its manager fonts, so a fixed plugin reloads new code.
- Resource limits are validated as non-negative numbers: 400 at
  POST /plugins/limits, bad cached records ignored with one warning.
  Route docstrings note health/metrics reset and limits only change the
  web process's view.
- SchemaManager.get_schema_path resolves each search dir via
  resolve_plugin_dir (manifest id, ledmatrix-<id>) before the literal
  paths; plugins/ still before plugin-repos/. Misses cached 30s and
  logged once at DEBUG.
- install_from_url sets an existing copy aside and restores it if the
  move fails, under the per-plugin reinstall lock.
- Operation queue refuses a second pending op for a plugin and trims
  _operations with history.
- get_vegas_render_width reads display_manager.width first.
- get_logger in store/schema/health/resource/saved_repositories;
  UTF-8 reads in store_manager and state_manager.
- Docs: update_interval precedence (manifest over config) stated where
  users are told to set it in config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): build the limits 400 message from the field name, not an exception

CodeQL flagged str(e) flowing into the response. invalid_limit_field()
returns the offending field without raising, and limits_from_dict uses it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-28 10:41:40 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent 0e9e2cabba
commit c00bf5e8e6
24 changed files with 869 additions and 49 deletions
+18
View File
@@ -85,6 +85,17 @@ class PluginOperationQueue:
f"Plugin {plugin_id} already has an active operation: "
f"{active_op.operation_id} ({active_op.operation_type.value})"
)
# _active_operations only holds the *running* one, so a second
# request while the first still waits in the queue (a double-
# clicked Install) used to be queued too, and both ran back to
# back. Refuse it the same way.
for queued_op in self._operations.values():
if queued_op.plugin_id == plugin_id and queued_op.status == OperationStatus.PENDING:
raise ValueError(
f"Plugin {plugin_id} already has an active operation: "
f"{queued_op.operation_id} ({queued_op.operation_type.value})"
)
# Create operation
operation = PluginOperation(
@@ -288,7 +299,14 @@ class PluginOperationQueue:
if len(self._operation_history) > self.max_history:
# Remove oldest operations
self._operation_history.sort(key=lambda op: op.created_at)
dropped = self._operation_history[:-self.max_history]
self._operation_history = self._operation_history[-self.max_history:]
# ...and forget them in the status map too, which otherwise kept
# every operation ever enqueued for the life of the process. A
# still-pending or running one is never dropped from lookups.
for op in dropped:
if op.status not in (OperationStatus.PENDING, OperationStatus.RUNNING):
self._operations.pop(op.operation_id, None)
def shutdown(self) -> None:
"""Shutdown the operation queue and worker thread."""