feat(ipc)!: remove the cache-key mailboxes (control socket stage 5)

The control socket is now the only way the web interface sends the display a
command. The display stops reading display_on_demand_request and
plugin_error_clear_request, and the web interface stops writing them.

- Display: no mailbox poll (MailboxWatch, the 1 s / 0.25 s cadence,
  _consume_on_demand_request, the deprecation log) and no persisted
  display_on_demand_processed_id guard; the error publisher reads no clear
  request. CacheManager.file_signature and MailboxWatch are removed.
- A write to either retired key is dropped by CacheManager.save_cache and
  logged once per writer, naming the plugin from the call stack (or the
  request's plugin_id), with the API to move to.
- Web: on-demand start with no display listening starts the service (when
  start_service) and sends the request again once the socket answers (45 s,
  10 s for a running service without a socket yet); every other failure is
  a 503 (400 for invalid_args). Stop answers 503 when no display listens,
  unless stop_service. errors/clear answers 503 with a reason-specific
  message instead of writing a request; clear_pending is always false.
  src.ipc.client.should_fall_back is replaced by display_not_listening.
- Kept: display_current_state, display_on_demand_state,
  plugin_runtime_snapshot and the heartbeat (read whenever the socket cannot
  answer), and display_on_demand_config (the display's resume record).

Tests: mailbox-only tests removed (test_on_demand_mailbox.py, the mailbox
cadence, file_signature and MailboxWatch tests); tests that injected
requests through the mailbox now use the socket queue or a plugin's
in-process request. The run-loop harness sends on-demand requests over its
fake control socket, so four golden traces change: on-demand starts and
stops land at the request instant instead of the next 0.25 s mailbox look
(one frame fewer on the screen they end), and in vegas.json within one
frame instead of 263 ms, which shifts the later 1 s-throttled WiFi-notice
check by under a second.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-10-05 09:32:22 -04:00
co-authored by Claude Opus 5.5
parent bb475a79ea
commit c461f4efdb
33 changed files with 1307 additions and 1756 deletions
+16 -43
View File
@@ -650,11 +650,10 @@ When nothing is running on demand, `data.state` is
> above (or the web UI buttons). The API handlers
> (`start_on_demand_display()` / `stop_on_demand_display()` in
> `web_interface/blueprints/api_v3/display.py`) send the request over the
> display's control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)).
> Only when the socket cannot carry it do they write it into the cache
> manager under the `display_on_demand_request` key, which
> `DisplayController._poll_on_demand_requests()`
> (`src/display_controller.py`) picks up. A separate
> display's control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)),
> the only way in (the `display_on_demand_request` cache-key mailbox is
> gone). A plugin asks for the screen with `BasePlugin.request_on_demand()`
> (see [PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md)). A separate
> `display_on_demand_config` key is used by the controller itself
> during activation (`_activate_on_demand()`) to track what's
> currently running, and is cleared by `_clear_on_demand()`.
@@ -737,29 +736,12 @@ keys helps troubleshoot stuck states.
### Cache Keys
**1. display_on_demand_request** (TTL: 1 hour)
```json
{
"request_id": "uuid-string",
"action": "start|stop",
"plugin_id": "plugin-name",
"mode": "mode-name",
"duration": 30.0,
"pinned": true,
"timestamp": 1234567890.123
}
```
**Purpose:** Communication from web interface to display controller, as the
fallback when the control socket cannot carry the request (deprecated; it
will be removed in a later release)
**When Set:** API endpoint receives a request and the display's control
socket is unavailable (display stopped, or older than the socket or the
command); some plugins also write it directly
**Read:** once a second while the display serves the control socket (0.25 s
without it), and only when the file changed since the last look
**Auto-Cleared:** After processing or 1 hour TTL
Requests are not cache keys: they go over the control socket. The
`display_on_demand_request` and `display_on_demand_processed_id` keys of
earlier releases are no longer written or read; a leftover file is
harmless and can be deleted.
**2. display_on_demand_config** (No TTL)
**1. display_on_demand_config** (No TTL)
```json
{
"mode": "mode-name",
@@ -771,7 +753,7 @@ without it), and only when the file changed since the last look
**When Set:** Controller processes start request
**Auto-Cleared:** When on-demand stops
**3. display_on_demand_state** (Continuously updated)
**2. display_on_demand_state** (Continuously updated)
```json
{
"active": true,
@@ -785,31 +767,24 @@ without it), and only when the file changed since the last look
**When Set:** Every display loop iteration
**Auto-Cleared:** Never (continuously updated)
**4. display_on_demand_processed_id** (TTL: 1 hour)
```text
"uuid-string-of-last-processed-request"
```
**Purpose:** Prevents duplicate request processing
**When Set:** After processing request
**Auto-Cleared:** After 1 hour TTL
### When Manual Clearing is Needed
**Scenario 1: Stuck in On-Demand State**
- Symptom: Display stays on one plugin, won't return to rotation
- Clear: `config`, `state`, `request`
- Clear: `config`, `state`
**Scenario 2: Mode Switching Issues**
- Symptom: Can't change to different plugin
- Clear: `request`, `processed_id`, `state`
- Clear: `state`, then restart the display
**Scenario 3: On-Demand Not Activating**
- Symptom: Button click does nothing
- Clear: `processed_id`, `request`
- Symptom: Button click does nothing, or answers an error
- Check the error's `socket_error` (see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md),
"Checking it on a device"); no cache key is involved
**Scenario 4: After Service Crash**
- Symptom: Strange behavior after crash/restart
- Clear: All four keys
- Clear: both keys
### Manual Recovery Procedures
@@ -846,8 +821,6 @@ from src.cache_manager import CacheManager
cache = CacheManager()
cache.clear_cache('display_on_demand_config')
cache.clear_cache('display_on_demand_state')
cache.clear_cache('display_on_demand_request')
cache.clear_cache('display_on_demand_processed_id')
```
> `CacheManager` also has a `delete(key)` method — a thin wrapper over
+12 -15
View File
@@ -42,11 +42,11 @@ each other. They share three things:
| State | Where | Written by | Read by |
|---|---|---|---|
| On-demand command | control socket `/run/ledmatrix/control.sock` ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py), via [`src/ipc/client.py`](../src/ipc/client.py) | display: [`src/ipc/server.py`](../src/ipc/server.py) acks; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand request (fallback) | cache `display_on_demand_request` | web, only when the socket could not carry the request (`should_fall_back`); four plugins write it directly | display: `_poll_on_demand_requests()`, a `stat()` every 1 s while the socket is up (0.25 s without), read only when the file changed |
| On-demand request from a plugin | in memory: `BasePlugin.request_on_demand()` / `end_on_demand()` | a plugin in the display process | display: `submit_plugin_on_demand()` queues it; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand state | cache `display_on_demand_state` | display: `_publish_on_demand_state()` | web: `/api/v3/display/on-demand/status` |
| Current screen | cache `display_current_state` | display | web: `/api/v3/display/current-status` |
| Plugin errors | cache `plugin_error_snapshot` | display: `ErrorSnapshotPublisher` ([`src/error_aggregator.py`](../src/error_aggregator.py)) | web: `read_error_report()` for `/api/v3/errors/*` |
| Error clear | control socket `errors.clear`; cache `plugin_error_clear_request` as the fallback | web: `POST /api/v3/errors/clear` | display: applied before the socket answers; the mailbox on the error publisher's 5 s tick, read only when the file changed |
| Error clear | control socket `errors.clear` | web: `POST /api/v3/errors/clear` | display: applied before the socket answers |
| Font usage | cache `font_usage_snapshot` | display: `FontUsagePublisher` ([`src/font_usage.py`](../src/font_usage.py)) | web: Fonts tab |
| Fetch statistics (requests per plugin and host) | cache `fetch_stats_snapshot` | display: `FetchStatsPublisher` ([`src/common/fetch_service.py`](../src/common/fetch_service.py)), at most once a minute on change | web: `read_fetch_stats()` for `/api/v3/plugins/fetch-stats` |
| Plugin health | cache `plugin_health:<id>` | display (web writes on reset) | web: `/api/v3/plugins/health` |
@@ -58,18 +58,15 @@ each other. They share three things:
The on-demand start route starts `ledmatrix.service` when it is not running
(`start_service`, on by default) but never restarts a running one. The routes
send the command over the display's control socket and get an ack; only when
the socket could not carry it (a stopped display, one older than the socket
or the command) do they write the mailbox instead. A display that had the
request and refused it is answered with the error, not posted a mailbox
copy. The display looks at the mailbox every
`MAILBOX_POLL_INTERVAL_WITH_SOCKET` (1 s) while it serves the socket, and
every `ON_DEMAND_POLL_INTERVAL` (0.25 s) without one, from its dwell sleep,
its render loops and Vegas's interrupt check as well as the main loop; a
look is one `stat()` unless the file changed. Both ways end in the same
handler, `_handle_on_demand_request()`.
send the command over the display's control socket and get an ack. That is
the only way in: the cache-key mailboxes (`display_on_demand_request`,
`plugin_error_clear_request`) are gone, and a write to either is dropped
with a warning. When no display is listening yet, the start route starts the
service (if asked) and sends the request again once the socket is up; any
other failure is answered as an error. Socket commands and plugins'
in-process requests end in the same handler, `_handle_on_demand_request()`.
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
for the protocol, the permission model and the plan to retire the mailboxes.
for the protocol and the permission model.
### Web and display processes: who runs plugins
@@ -118,8 +115,8 @@ other web-UI action runs its script as a subprocess. A later, explicit
The **control socket** from the web process to the display
([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) carries on-demand
commands and reloads an updated plugin; its next stages stream the
display's state and retire the cache-key mailboxes. The plugin web-entry
commands, reloads an updated plugin and streams the display's state; it
replaced the cache-key mailboxes. The plugin web-entry
contract above is still to come.
### Plugin state: desired, observed, and who owns it
+90 -89
View File
@@ -1,18 +1,18 @@
# Control socket (web → display)
The display process serves a Unix socket that the web interface uses to send
it commands and get an answer back. It replaces the cache-file "mailboxes" on
it commands and get an answer back. It replaced the cache-file "mailboxes" on
the SD card one command at a time. Stage 1 carries on-demand start, stop and
status. Stage 2 makes those commands land within a frame on every kind of
screen, and adds `brightness.set` and `plugin.reload`. Stage 3 adds a state
stream (`state.get`, `state.subscribe`), so the web interface reads what the
display is doing from the socket instead of from cache files the display
wrote to the SD card. Stage 4 makes the socket the only way a command goes
while it works: the web interface writes a mailbox only when the socket
cannot carry the request, `errors.clear` replaces the last command that
always went through a mailbox, and the display looks at the mailboxes once
a second, with a `stat()`. The file mailboxes and the cache keys stay as a
fallback for one release.
display wrote to the SD card. Stage 4 makes the socket the only way a command goes
while it works, and `errors.clear` replaces the last command that always
went through a mailbox. Stage 5 removes the mailboxes: the web interface no
longer writes them and the display no longer reads them, so the socket is
the only way a command reaches the display (see "Without the socket"). The
cache keys the display writes for readers stay as their fallback.
| | |
|---|---|
@@ -145,14 +145,17 @@ than the display's wait, so `pending` arrives before the client gives up.
`v` the display does not speak gets `unsupported_version`. `hello` is checked
by its `versions` list instead, and its result names the highest version both
sides share, so a client can find out what a display supports before it
relies on anything newer. The client sends `v: 1` and falls back to the
mailbox when the display refuses it. It does not send `hello` first, which
relies on anything newer. The client sends `v: 1`, and a route answers an
error when the display refuses it. It does not send `hello` first, which
saves a round trip.
New commands are added within a version, so stage 2 is still version 1. A
display that does not know a command answers `unknown_command`, which the
web interface treats like any other socket failure and falls back from, and
`hello` lists the commands a display knows. The version changes only when the
`hello` lists the commands a display knows.
(Since stage 5, a command the display does not know is an error the route
reports, telling the user to restart the display; there is no mailbox left
to fall back to.) The version changes only when the
envelope or the meaning of an existing command changes.
**Events.** `state.subscribe` is the one command with more than one message
@@ -368,13 +371,12 @@ request, validates it against the contract, and then does one of two things:
- For a query, it answers from a status snapshot the display provides
(`DisplayController._control_status`). The snapshot only reads attributes.
The render thread drains the queue in `_poll_on_demand_requests()`, the same
place it reads the mailbox:
The render thread drains the queue in `_poll_on_demand_requests()`:
- An on-demand command goes to `_handle_on_demand_request()`, which is the
mailbox's own handler. The two paths share all of their code: activation,
the processed-id guard, error publishing, and resuming the rotation
afterwards.
- An on-demand command goes to `_handle_on_demand_request()`, which also
handles plugins' own requests. The two paths share all of their code:
activation, the request-id guard, error publishing, and resuming the
rotation afterwards.
- `brightness.set` is applied there and then (`_apply_control_brightness`),
and the current frame is pushed again so the panel shows it.
- `plugin.reload` starts at the top of the next loop pass, the place where
@@ -407,14 +409,13 @@ place it reads the mailbox:
unloads it mid-load; a disable saved meanwhile is applied once the
reload is done. A second reload of the same plugin runs after the first.
The floor on the mailbox read (0.25 s, 1 s since stage 4) does not apply to the queue, because
draining it costs no disk read. A queued command also lets
`_service_pending_changes()` skip its own floor.
Draining the queue costs no disk read, so it has no floor. A queued command
also lets `_service_pending_changes()` skip its own 0.25 s floor.
### Waking the render thread (stage 2)
Stage 1 made the socket answer, but not land sooner: a queued command waited
for the same polls the mailbox does. Measured on ledpi (Pi 4, 24 fps Vegas),
for the same polls the mailbox did. Measured on ledpi (Pi 4, 24 fps Vegas),
a start took 1.02 s on a static screen and about 0.4 s in Vegas either way.
Now the queue wakes the render thread:
@@ -434,8 +435,7 @@ Now the queue wakes the render thread:
- **Scrolling screens** already service pending changes every frame.
So a command lands within a millisecond or so on a static screen and in a
dwell, and within one frame in Vegas and on a scrolling screen. The mailbox
is slower on purpose (see "The mailboxes now"). Commands still run only on the render thread: the
dwell, and within one frame in Vegas and on a scrolling screen. Commands still run only on the render thread: the
connection threads only queue them and set the event. The one exception is
the slow half of `plugin.reload` (tearing down and loading the plugin),
which runs on its own thread. Every change to the display's state still
@@ -453,74 +453,75 @@ bookkeeping. A client's send to the render thread waking took 0.72 ms median
Without a socket (Windows, `LEDMATRIX_CONTROL_SOCKET=off`) the waits are the
plain sleeps they were.
**Exactly once.** Since stage 4 the web interface writes the mailbox only
when the display never had the request (see "When the web interface falls
back"), so a request goes one way or the other, never both. A command and a
mailbox write for the same request still share one `request_id`, and the
`on_demand_request_id` and processed-id checks still drop a second copy: an
older web interface (before stage 4) wrote the mailbox after a reply timed
out, too. The display takes such a copy out of the mailbox when it drops it.
**Exactly once.** A request goes over the socket once. The web route sends
it again only while no display is listening (see "Without the socket"), so
a display gets it at most once; a start with a request id the display has
just processed is dropped anyway (`on_demand_request_id`). The persisted
`display_on_demand_processed_id` guard against a mailbox replayed after a
restart went with the mailbox.
## When the web interface falls back (stage 4)
## Without the socket (stage 5)
The client tells a request the display never had from one it had and then
failed. `ControlError.sent` is True once the whole request was written to a
connected display; a refusal the display sends before reading anything
(`forbidden`, too many connections) carries no request id, and leaves it
False. `src.ipc.client.should_fall_back()` is the one rule every route uses:
False. `src.ipc.client.display_not_listening()` picks out the one case a
later retry can fix: nothing is listening (`no_socket`, `refused`) and the
request was never sent.
| What happened | Example reasons | Mailbox? | The route answers |
|---|---|---|---|
| The display never had it | `no_socket`, `refused`, `disabled`, `unsupported`, a connect or send that timed out, `forbidden` / `busy` at the door, `invalid_request` (refused by the client itself) | yes | success, `transport: "mailbox"`, `socket_error` |
| A display too old to know it (the upgrade case) | `unknown_command`, `unsupported_version` | yes | as above |
| The display had it and failed | `busy` (queue full), `invalid_args`, `internal`, a timeout or hang-up after the send, `bad_response` | no | `503` (`400` for `invalid_args`), `socket_error` |
| What happened | Example reasons | On-demand start | On-demand stop | `errors.clear` |
|---|---|---|---|---|
| No display listening | `no_socket`, `refused` | service stopped: `400` without `start_service`; with it, start the service and send again once the socket answers (up to 45 s), else `503`. Service running (still starting): send again for up to 10 s, else `503` | `503` ("not running" / "may still be starting"); with `stop_service` the service is stopped and the route succeeds | `503` ("not running"; its errors are the last run's, and the next run starts with none) |
| A display too old to know the command | `unknown_command`, `unsupported_version` | `503` | `503` | `503`, "restart it" |
| No socket in this process | `disabled`, `unsupported` (Windows, `LEDMATRIX_CONTROL_SOCKET=off`) | `503` | `503` | `503` |
| The display had it and failed, turned it away, or never answered | `busy`, `invalid_args`, `internal`, a timeout, a hang-up, `bad_response`, `forbidden` | `503` (`400` for `invalid_args`) | `503` (with `stop_service`: stopped anyway) | `503` |
A display that had the request may have applied it (a reply that timed out),
or would refuse the mailbox copy as well (bad arguments), or is stuck and
would not read the mailbox either (a full queue). Writing the copy anyway
only turned that into a "success". An on-demand stop with `stop_service`
still stops the service, which ends on-demand whatever happened.
Every error answer carries `socket_error` (a reason code, or `other`).
Nothing is written to the cache in any of these cases. The start route's
waits are bounded (`ON_DEMAND_SOCKET_WAIT_SECONDS`,
`ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS` in
`web_interface/blueprints/api_v3/display.py`): the socket comes up when the
display's run loop starts, after every plugin has loaded. A client with a
shorter HTTP timeout (the MQTT bridge's is 15 s) can give up first while
the route still delivers the request.
Brightness and plugin reload never had a mailbox: without the socket, the
config watcher applies the saved brightness and a reload becomes the
restart banner, as before.
### The mailboxes now
### The mailboxes are gone
| Mailbox | Written by | Read by the display | While the socket is up |
|---|---|---|---|
| `display_on_demand_request` | the web interface, only on fallback; plugins that predate `BasePlugin.request_on_demand()`, or run on a core without it | the render thread, `_poll_on_demand_requests()` | looked at every 1 s (`MAILBOX_POLL_INTERVAL_WITH_SOCKET`), 0.25 s without a socket |
| `plugin_error_clear_request` | the web interface, only on fallback | the error publisher's thread, every 5 s tick | unchanged rate |
| Former mailbox | Last written by | Now |
|---|---|---|
| `display_on_demand_request` | the web interface on fallback (stage 4); plugins that predate `BasePlugin.request_on_demand()` | not read. A write is dropped by `CacheManager.save_cache` (`RETIRED_MAILBOX_KEYS`) and logged once per writer as a warning that names the plugin when it can be told (the plugin instance on the call stack, else the `plugin_id` in the request) |
| `plugin_error_clear_request` | the web interface on fallback (stage 4) | not read; a write is dropped and logged the same way |
A look is one `stat()` of the mailbox file (`CacheManager.file_signature`):
`(inode, mtime, size)`, and every write renames a new file into place, so a
new write always looks different. `MailboxWatch` reads the file only when
that changed since the last look, so a mailbox that holds nothing new, or
nothing at all, costs no open and no parse. A socket command never reads or
deletes the on-demand mailbox. A start already processed is taken out of
the mailbox instead of being re-read until it expires.
A file left on the SD card by an older version is never read again, so it
is harmless. `MailboxWatch` and
`CacheManager.file_signature`, which made a look at a mailbox one `stat()`,
went with them.
A request that comes through the on-demand mailbox while the socket is up
is logged once per writer (`came through the file mailbox although the
control socket is up`), which names the plugins that still write it.
The four plugins that wrote `display_on_demand_request` (birdnet-go,
mqtt-notifications, on-air, pomodoro-timer) use `request_on_demand()` on a
core that has it, and fall back to the mailbox only when that method is
missing or answers `None` (no display in the process, or a full queue).
On a stage-5 core such a fallback write is dropped with the warning above.
### Plugins in the display process
A plugin asks for the screen with `BasePlugin.request_on_demand()` and gives
it back with `end_on_demand()` (see "On-demand display" in
[PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md)). Neither goes through
the socket or a file: `PluginManager` hands the mailbox-shaped request,
the socket or a file: `PluginManager` hands the on-demand request,
marked `source: 'plugin'`, to `DisplayController.submit_plugin_on_demand`,
which queues it in memory (at most `PLUGIN_ON_DEMAND_QUEUE_SIZE`, 32) from
whatever thread the plugin called on, and wakes the render thread through
the socket's queue flag (`ControlServer.wake()`). The render thread applies
it in `_drain_control_commands`, after the socket's commands, through the
same `_handle_on_demand_request`, so it lands within a frame like a socket
command. Without a socket it lands on the next pending-changes pass (typically
within 0.25 s). A plugin's stop ends only a session that plugin owns. The four
plugins that wrote the mailbox (birdnet-go, mqtt-notifications, on-air,
pomodoro-timer) use it where the core has it and write the mailbox
otherwise.
command. Without a socket it lands on the next pending-changes pass. A
plugin's stop ends only a session that plugin owns.
## Robustness
@@ -538,8 +539,7 @@ block the render loop or crash it:
found. A client that disconnects mid-message is dropped silently. No
exception from a handler leaves the connection thread.
- **Full queue.** When the queue is full, the client gets `busy`, and the web
interface answers `503` rather than write the mailbox, which the stuck
render thread would not read either. A full queue means the render thread
interface answers `503`. A full queue means the render thread
is stuck, and the systemd watchdog deals with that.
- **Awaited commands.** The wait for an awaited command's outcome happens on
its connection thread and is bounded (`AWAIT_SECONDS`), so a stuck render
@@ -554,8 +554,8 @@ block the render loop or crash it:
process created.
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
as before. The web interface then uses the mailbox, and reads the cache
keys and the heartbeat file.
as before. The web interface then cannot send it commands (the routes
answer `503`), and reads the cache keys and the heartbeat file.
- **Subscribers (stage 3).** A `state.subscribe` connection gives its request
slot back and takes one of 4 subscriber slots (`MAX_SUBSCRIBERS`). A fifth
gets `busy`. So a few browsers' web processes holding streams can never
@@ -613,11 +613,9 @@ device never touches the live display.
## Stage plan
1. **On-demand, with acks (done, #706).** Contract, server, client.
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes try the
socket first and report `transport: "socket" | "mailbox"` (plus
`socket_error` on fallback). The mailbox is unchanged, and the plugins that
write it directly (birdnet-go, mqtt-notifications, on-air, pomodoro-timer)
keep working.
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes tried
the socket first and reported `transport: "socket" | "mailbox"` (plus
`socket_error` on fallback).
2. **Commands that were restarts or polls (done).**
- The render thread waits on the queue instead of sleeping, and Vegas
checks it every frame, so a command lands within a frame on every kind
@@ -660,12 +658,12 @@ device never touches the live display.
4. **The mailboxes become a fallback (done).**
- The web interface writes a mailbox only when the socket could not carry
the request (`should_fall_back`); a display that had it and failed is
answered as that (see "When the web interface falls back").
answered as that.
- `errors.clear` replaces `plugin_error_clear_request` as the way a clear
reaches the display.
- The display looks at the on-demand mailbox once a second while the
socket is up, reads either mailbox only when its file changed, and logs
who still writes the on-demand one (see "The mailboxes now").
who still writes the on-demand one.
- Not changed, deliberately: config saves (the schedule, the dim
schedule, plugin settings) still reach the display through
`config.json` and its watcher, which is the setting itself rather than
@@ -674,15 +672,19 @@ device never touches the live display.
already stats at most once a second. Plugin health and metrics resets
write the persisted record the display publishes and do not reach the
running display (their routes say so); they are not mailboxes.
5. **Remove the mailboxes (next release).** Once every device has run a
display with stage 4, the web interface stops writing both mailboxes and
the display stops reading them. The four plugins that wrote
`display_on_demand_request` now have an in-process way to ask for the
screen (`BasePlugin.request_on_demand()` / `end_on_demand()`, see
"Plugins in the display process"); they keep the mailbox write only as
their fallback on older cores. The display also stops writing `display_current_state`,
`display_on_demand_state` and `plugin_runtime_snapshot` once the web
interface no longer falls back to them.
5. **Remove the mailboxes (done, the release after 3.8.1).** The web
interface no longer writes `display_on_demand_request` or
`plugin_error_clear_request`, and the display no longer reads them (see
"Without the socket"). When no display is listening, the start route
starts the service if asked and sends the request again once its socket
is up; every other failure is an error the route reports. A write to
either key is dropped with a one-time warning naming the writer. The
display still writes `display_current_state`, `display_on_demand_state`
and `plugin_runtime_snapshot`: the web interface reads them whenever the
socket cannot answer (a stopped or starting display, a web user not yet
in the socket's group, a platform without Unix sockets), and
`display_on_demand_config` is the display's own record for resuming a
session after a restart. Retiring those keys is left for later.
## Checking it on a device
@@ -694,12 +696,11 @@ curl -s -X POST localhost:5000/api/v3/display/on-demand/start \
# ... "transport": "socket"
```
If the response says `"transport": "mailbox"`, `socket_error` gives the
reason. `no_socket` means the display is stopped or predates the socket.
`refused` usually means the web user is not in the socket's group, which
takes effect when the web service restarts after the user is added. A `503`
with `"transport": "socket"` means the display had the request and did not
take it (`busy`, `timeout`, ...): nothing was written to the mailbox.
On an error, `socket_error` gives the reason. `no_socket` means the display
is stopped or still starting. `refused` usually means the web user is not in
the socket's group, which takes effect when the web service restarts after
the user is added. `busy`, `timeout` and the like mean the display had the
request and did not take it. Nothing is ever written to a mailbox.
An error clear:
@@ -707,7 +708,7 @@ An error clear:
curl -s -X POST localhost:5000/api/v3/errors/clear \
-H 'Content-Type: application/json' -d '{"all":true}'
# ... "applied": true, "transport": "socket"
sudo journalctl -u ledmatrix | grep -E "Cleared .* plugin error|file mailbox"
sudo journalctl -u ledmatrix | grep -E "Cleared .* plugin error|retired"
```
Brightness and a plugin reload:
+30 -20
View File
@@ -522,42 +522,52 @@ Returns the request id once queued, or `None` as above.
These methods are new after core 3.8.0 (see `CHANGELOG.md`). Before them,
plugins wrote the `display_on_demand_request` cache key (the "mailbox")
themselves. The display reads it only once a second while the control
socket is up, and it will be removed in a future release (see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md), stage 5). A plugin that
must keep working on older cores checks for the method, and writes the
mailbox only when the method is missing or answers `None`:
themselves. **The mailbox is gone** (see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md), stage 5): the display no
longer reads it, and a write to it is dropped with a warning in the log,
once per plugin:
```
Ignored a write to the retired 'display_on_demand_request' cache key by plugin 'my-plugin': ...
```
A plugin that only needs to run on cores with these methods calls them and
treats `None` as "no display took it":
```python
def _show_alert(self):
if self.request_on_demand(mode="my_alert", duration=15) is None:
self.logger.info("No display to show the alert on")
```
A plugin that must also work on cores before them can keep the mailbox
write as its fallback, guarded by `hasattr`: on those cores the display
still reads it, and on a current core the write is only dropped and logged
(when the method is missing it never runs at all). Raise
`ledmatrix_min_version` to the release that added the methods once you no
longer need it.
```python
import time, uuid
def _show_alert(self):
if hasattr(self, "request_on_demand") and self.request_on_demand(
mode="my_alert", duration=15):
if hasattr(self, "request_on_demand"):
self.request_on_demand(mode="my_alert", duration=15)
return
# Older core, or no display in this process: the mailbox, as before.
# A core older than request_on_demand(): the mailbox it still reads.
self.cache_manager.set("display_on_demand_request", {
"request_id": str(uuid.uuid4()), "action": "start",
"plugin_id": self.plugin_id, "mode": "my_alert",
"duration": 15, "pinned": False, "timestamp": time.time(),
})
def _release(self):
if hasattr(self, "end_on_demand") and self.end_on_demand():
return
self.cache_manager.set("display_on_demand_request", {
"request_id": str(uuid.uuid4()), "action": "stop",
"plugin_id": self.plugin_id, "timestamp": time.time(),
})
```
Keep `ledmatrix_min_version` where it is: the fallback is what keeps the
plugin working on older cores. A mailbox stop ends any on-demand session,
whoever started it; `end_on_demand()` ends only the plugin's own.
`end_on_demand()` ends only the plugin's own session; a stop written to the
mailbox on an older core ends any session, whoever started it.
Both methods answer a request id only when the plugin manager returned a
string, so a test that gives the plugin a `MagicMock()` plugin manager gets
`None` and exercises the mailbox path. To test the new path, set
`None`. To test the path where the display takes the request, set
`plugin_manager.request_on_demand.return_value = "some-id"`.
> The full source for `BasePlugin` lives in
+36 -34
View File
@@ -464,7 +464,7 @@ Request a specific plugin to display on-demand.
- `mode` (string, optional): Display mode name (plugin_id inferred if not provided)
- `duration` (number, optional): Duration in seconds (0 = until stopped)
- `pinned` (boolean, optional): Pin display (pause rotation)
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket (within about a second through the mailbox fallback). When false and the service is stopped, the route returns 400.
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket. A stopped one is started and sent the request once its socket is up, which can take as long as the display takes to load its plugins (the route waits up to 45 s). When false and the service is stopped, the route returns 400.
**Response**:
```json
@@ -484,24 +484,30 @@ Request a specific plugin to display on-demand.
`service` is `null` when `start_service` is false.
`transport` says how the request reached the display: `"socket"` means the
display's control socket acknowledged it (it is queued for the render thread,
which wakes for it and applies it within a frame; see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), `"mailbox"` means it was
written to the cache mailbox the display polls, as before the socket existed.
The mailbox is used only when the socket could not carry the request. With
`"mailbox"`, `socket_error` gives the reason (`no_socket` when the display is
stopped or predates the socket, `refused`, a connect `timeout`,
`unknown_command` from a display too old for the command, ...). Either way
the request is applied the same way; `request_id` is the same id in both.
`transport` is always `"socket"`: the display's control socket acknowledged
the request (it is queued for the render thread, which wakes for it and
applies it within a frame; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)).
The `"mailbox"` value earlier releases could answer is gone with the
mailbox: nothing is written to the cache.
When the display had the request and did not take it -- a full queue
(`busy`), bad arguments (`invalid_args`), no answer after the request was
sent (`timeout`, `closed`) -- the route answers `503` (`400` for
`invalid_args`) with `status: "error"` and `data: {request_id, transport:
"socket", socket_error}`, and writes nothing to the mailbox. The stop route
does the same, except that with `stop_service: true` it still stops the
service and answers success.
When the display did not take the request, the route answers an error with
`status: "error"` and `data: {request_id, transport: "socket",
socket_error}` (plus `service` when it started or checked the service):
- no display listening (`no_socket`, `refused`): with the service stopped
and `start_service` false, `400`; otherwise the route waits for the
display's socket (45 s after starting the service, 10 s when it was
already running and may still be starting) and answers `503` if it never
answers;
- a full queue (`busy`), no answer after the request was sent (`timeout`,
`closed`), a display older than the command (`unknown_command`), no
socket in the web process (`disabled`, `unsupported`): `503` at once;
- bad arguments (`invalid_args`): `400`.
The stop route answers the same errors (`503` when no display is
listening, with a message saying whether the service is stopped), except
that with `stop_service: true` it still stops the service and answers
success, with `socket_error` set.
### Stop On-Demand Display
@@ -531,7 +537,7 @@ Stop the current on-demand display.
}
```
`transport` and `socket_error` are as for start.
`transport` and the errors are as for start; a stop is never retried.
---
@@ -2197,7 +2203,7 @@ Every response below adds three fields to the shape it always had:
|---|---|
| `snapshot_available` | `false` until the display service has reported (for example, it is not running). Counts are then zero. |
| `generated_at` | When the display service produced the snapshot (ISO, the Pi's local time), or `null`. |
| `clear_pending` | A clear has been requested and the display service has not applied it yet. |
| `clear_pending` | Always `false`: a clear is applied before its route answers. Kept for compatibility. |
### Get Error Summary
@@ -2278,21 +2284,17 @@ before it answers: `applied` is `true`, `transport` is `"socket"`, and
}
```
When the socket cannot carry it (the display is stopped, or older than
`errors.clear`) the clear is asynchronous, as before: the web interface
records a request (`plugin_error_clear_request` in the shared cache),
`applied` is `false` and `transport` is `"mailbox"`, and the display service
applies it within about 5 seconds. Reads hide the cleared errors from the
moment the request is recorded. Until the display service applies an
age-based clear, `recent_errors` and `active_patterns` are already filtered
but the counts are the old ones, and `clear_pending` is `true`. Then
`cleared_count` is how many of the reported errors the clear hides, and
`null` when that cannot be known before the display service applies it (an
age-based clear over more errors than the report lists).
`clear_requested`, `applied` and `transport` are always `true`, `true` and
`"socket"`, and are kept for compatibility.
A request that could not be written to the shared cache answers `500`. A
display that had the request and failed it (`internal`, a timeout after the
request was sent) answers `503`, with `context.socket_error`.
When the socket cannot carry the clear, the route answers `503` with
`context.socket_error`, and nothing is cleared or recorded: the
`plugin_error_clear_request` mailbox earlier releases fell back to is gone.
The message says why: the display service is not running (its errors are
then the last run's, and its next run starts with none), it is too old for
`errors.clear` (restart it), the web process has no socket, or the display
had the request and failed it (`internal`, a timeout after the request was
sent).
---