Compare commits

..
Author SHA1 Message Date
ChuckandClaude Opus 5.5 8879886e50 fix(web): keep exception messages out of API responses (py/stack-trace-exposure)
CodeQL had ~40 open py/stack-trace-exposure alerts on main. Almost all
flowed through describe_exception(), which returned "TypeName: message"
(redacted, capped); the rest through _run_systemctl_command's str(err),
WiFiManager's `return False, str(e)`, unit_refresh's f-strings and two
str(e)/f"{err}" messages in api_v3/__init__.py.

describe_exception() now returns a reason code -- the type, plus the
errno symbol for an OSError ("OSError:EIO", "PermissionError:EACCES") --
and logs the redacted message itself. That keeps what #538 wanted (a
failing disk still says EIO in the response) without quoting paths,
URLs or library internals, and fixes every call site at once; the
test_no_api_v3_handler_discards_its_exception policy still holds.

Service results: _get_display_service_status returns active/returncode
only, and the on-demand start/stop `service` result keeps
returncode/active/started/status but drops systemctl stdout/stderr
(logged on failure). Nothing in web_interface/static, the templates or
the MQTT bridge reads those fields. The Starlark SIGKILL-restart error
no longer returns systemctl stderr as `details`.

WiFi, unit-refresh, config-save and plugin-removal failures now say
what failed with the reason code and point at the log. display.py is
untouched (draft #773 edits it).

Tests: test_api_v3_no_exception_text.py drives one route per affected
file with a marker in the exception message and asserts it never
reaches the body; all 13 fail on origin/main, and targeted mutations
(drop the service filter, put stderr back, str(e) in WiFiManager,
{e} in unit_refresh, {install_err} in system.py, message back in
describe_exception) each fail at least one. Tests that asserted the old
message-in-details contract now assert the reason code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:41:07 -04:00
57 changed files with 2049 additions and 2398 deletions
+9 -56
View File
@@ -19,62 +19,15 @@ accepts both, but the store flags the old spelling as deprecated
## Unreleased
### Removed: the cache-key mailboxes (control socket stage 5) -- breaking
The control socket (`docs/IPC_CONTROL_SOCKET.md`) is now the only way the
web interface sends the display a command. The file mailboxes it fell back
to for one release are gone.
- **`display_on_demand_request` and `plugin_error_clear_request` are no
longer written or read.** The display stops polling the on-demand mailbox
(`MailboxWatch`, `MAILBOX_POLL_INTERVAL_WITH_SOCKET`,
`_consume_on_demand_request` and the persisted
`display_on_demand_processed_id` guard are removed), and the error
publisher stops reading the clear request. `CacheManager.file_signature`,
`src.cache_manager.MailboxWatch`, `src.ipc.client.should_fall_back` and
`src.error_aggregator.ERROR_CLEAR_REQUEST_KEY` are removed.
- **A write to either key is dropped, with one warning per writer.**
`CacheManager.save_cache` (and so `set`) refuses `RETIRED_MAILBOX_KEYS`
and logs `Ignored a write to the retired '<key>' cache key by plugin
'<id>'`, naming the plugin from the call stack or the request. Plugins
must use `BasePlugin.request_on_demand()` / `end_on_demand()` (3.8.1).
In the official monorepo, birdnet-go, mqtt-notifications, on-air and
pomodoro-timer still write the mailbox, but only as their fallback when
those methods are missing or answer `None`.
- **On-demand routes without a listening display.** `POST
/api/v3/display/on-demand/start` with no display listening starts the
service (when `start_service`, the default) and answers **`202`** with
`status: "starting"` at once; a single background worker in the web
process (`web_interface/on_demand_dispatch.py`) sends the request until
the display acknowledges it, for up to 45 s (10 s for a service that is
running but has no socket yet). `GET /display/on-demand/status` reports it
(`starting`; still `starting` with `delivered: true` after the display
acknowledges it, until the display publishes the state for that
`request_id`, now part of its on-demand state, for at most 30 s; then the
display's state, or `error` / `start-timeout`), and
`/display/current-status` adds `on_demand_pending`. A newer start replaces
a pending one and a stop cancels it (`cancelled_request_id`). The web UI
and the MQTT bridge treat `202` as taken. With the service stopped and
`start_service` false it answers `400`. Every other socket failure (`unknown_command`
from an older display, `disabled`/`unsupported`, `busy`, a timeout) is a
`503`. `/stop` answers `503` when no display is listening, unless
`stop_service` stops the service. `transport` is always `"socket"`; the
`"mailbox"` value is gone.
- **`POST /api/v3/errors/clear` without the socket answers `503`** (with
`context.socket_error` and a message saying why) instead of recording a
request. `clear_pending` in the error routes is now always `false`;
`src.error_aggregator.read_error_report()` returns only the snapshot, and
`error_summary_from_report()` / `plugin_health_from_report()` /
`request_error_clear()` lose their clear-request arguments.
- **Kept:** the display still writes `display_current_state`,
`display_on_demand_state`, `plugin_runtime_snapshot` and the heartbeat
file, because the web interface reads them whenever the socket cannot
answer (a stopped or starting display, a web user not yet in the socket's
group, Windows), and `display_on_demand_config`, its own record for
resuming a session after a restart.
- **Windows and `LEDMATRIX_CONTROL_SOCKET=off`:** with no socket, the web
interface can no longer start or stop on-demand sessions or clear errors
on a running display (the mailbox used to carry them).
- Web API error responses no longer carry an exception's message (CodeQL
`py/stack-trace-exposure`). `describe_exception()` now returns a reason
code -- the exception type, plus the errno for an `OSError`
(`OSError:EIO`, `PermissionError:EACCES`) -- and logs the message
instead, so `details` still names the fault without quoting paths, URLs
or library internals. The display service status and the on-demand
start/stop `service` results keep `active`, `returncode` and `started`
but drop systemctl's `stdout`/`stderr`; WiFi, unit-refresh and
config-save failures say what failed and point at the log.
## 3.8.2
+43 -16
View File
@@ -650,10 +650,11 @@ When nothing is running on demand, `data.state` is
> above (or the web UI buttons). The API handlers
> (`start_on_demand_display()` / `stop_on_demand_display()` in
> `web_interface/blueprints/api_v3/display.py`) send the request over the
> display's control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)),
> the only way in (the `display_on_demand_request` cache-key mailbox is
> gone). A plugin asks for the screen with `BasePlugin.request_on_demand()`
> (see [PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md)). A separate
> display's control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)).
> Only when the socket cannot carry it do they write it into the cache
> manager under the `display_on_demand_request` key, which
> `DisplayController._poll_on_demand_requests()`
> (`src/display_controller.py`) picks up. A separate
> `display_on_demand_config` key is used by the controller itself
> during activation (`_activate_on_demand()`) to track what's
> currently running, and is cleared by `_clear_on_demand()`.
@@ -736,12 +737,29 @@ keys helps troubleshoot stuck states.
### Cache Keys
Requests are not cache keys: they go over the control socket. The
`display_on_demand_request` and `display_on_demand_processed_id` keys of
earlier releases are no longer written or read; a leftover file is
harmless and can be deleted.
**1. display_on_demand_request** (TTL: 1 hour)
```json
{
"request_id": "uuid-string",
"action": "start|stop",
"plugin_id": "plugin-name",
"mode": "mode-name",
"duration": 30.0,
"pinned": true,
"timestamp": 1234567890.123
}
```
**Purpose:** Communication from web interface to display controller, as the
fallback when the control socket cannot carry the request (deprecated; it
will be removed in a later release)
**When Set:** API endpoint receives a request and the display's control
socket is unavailable (display stopped, or older than the socket or the
command); some plugins also write it directly
**Read:** once a second while the display serves the control socket (0.25 s
without it), and only when the file changed since the last look
**Auto-Cleared:** After processing or 1 hour TTL
**1. display_on_demand_config** (No TTL)
**2. display_on_demand_config** (No TTL)
```json
{
"mode": "mode-name",
@@ -753,7 +771,7 @@ harmless and can be deleted.
**When Set:** Controller processes start request
**Auto-Cleared:** When on-demand stops
**2. display_on_demand_state** (Continuously updated)
**3. display_on_demand_state** (Continuously updated)
```json
{
"active": true,
@@ -767,24 +785,31 @@ harmless and can be deleted.
**When Set:** Every display loop iteration
**Auto-Cleared:** Never (continuously updated)
**4. display_on_demand_processed_id** (TTL: 1 hour)
```text
"uuid-string-of-last-processed-request"
```
**Purpose:** Prevents duplicate request processing
**When Set:** After processing request
**Auto-Cleared:** After 1 hour TTL
### When Manual Clearing is Needed
**Scenario 1: Stuck in On-Demand State**
- Symptom: Display stays on one plugin, won't return to rotation
- Clear: `config`, `state`
- Clear: `config`, `state`, `request`
**Scenario 2: Mode Switching Issues**
- Symptom: Can't change to different plugin
- Clear: `state`, then restart the display
- Clear: `request`, `processed_id`, `state`
**Scenario 3: On-Demand Not Activating**
- Symptom: Button click does nothing, or answers an error
- Check the error's `socket_error` (see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md),
"Checking it on a device"); no cache key is involved
- Symptom: Button click does nothing
- Clear: `processed_id`, `request`
**Scenario 4: After Service Crash**
- Symptom: Strange behavior after crash/restart
- Clear: both keys
- Clear: All four keys
### Manual Recovery Procedures
@@ -821,6 +846,8 @@ from src.cache_manager import CacheManager
cache = CacheManager()
cache.clear_cache('display_on_demand_config')
cache.clear_cache('display_on_demand_state')
cache.clear_cache('display_on_demand_request')
cache.clear_cache('display_on_demand_processed_id')
```
> `CacheManager` also has a `delete(key)` method — a thin wrapper over
+15 -14
View File
@@ -42,11 +42,11 @@ each other. They share three things:
| State | Where | Written by | Read by |
|---|---|---|---|
| On-demand command | control socket `/run/ledmatrix/control.sock` ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py), via [`src/ipc/client.py`](../src/ipc/client.py) | display: [`src/ipc/server.py`](../src/ipc/server.py) acks; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand request from a plugin | in memory: `BasePlugin.request_on_demand()` / `end_on_demand()` | a plugin in the display process | display: `submit_plugin_on_demand()` queues it; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand request (fallback) | cache `display_on_demand_request` | web, only when the socket could not carry the request (`should_fall_back`); four plugins write it directly | display: `_poll_on_demand_requests()`, a `stat()` every 1 s while the socket is up (0.25 s without), read only when the file changed |
| On-demand state | cache `display_on_demand_state` | display: `_publish_on_demand_state()` | web: `/api/v3/display/on-demand/status` |
| Current screen | cache `display_current_state` | display | web: `/api/v3/display/current-status` |
| Plugin errors | cache `plugin_error_snapshot` | display: `ErrorSnapshotPublisher` ([`src/error_aggregator.py`](../src/error_aggregator.py)) | web: `read_error_report()` for `/api/v3/errors/*` |
| Error clear | control socket `errors.clear` | web: `POST /api/v3/errors/clear` | display: applied before the socket answers |
| Error clear | control socket `errors.clear`; cache `plugin_error_clear_request` as the fallback | web: `POST /api/v3/errors/clear` | display: applied before the socket answers; the mailbox on the error publisher's 5 s tick, read only when the file changed |
| Font usage | cache `font_usage_snapshot` | display: `FontUsagePublisher` ([`src/font_usage.py`](../src/font_usage.py)) | web: Fonts tab |
| Fetch statistics (requests per plugin and host) | cache `fetch_stats_snapshot` | display: `FetchStatsPublisher` ([`src/common/fetch_service.py`](../src/common/fetch_service.py)), at most once a minute on change | web: `read_fetch_stats()` for `/api/v3/plugins/fetch-stats` |
| Plugin health | cache `plugin_health:<id>` | display (web writes on reset) | web: `/api/v3/plugins/health` |
@@ -58,17 +58,18 @@ each other. They share three things:
The on-demand start route starts `ledmatrix.service` when it is not running
(`start_service`, on by default) but never restarts a running one. The routes
send the command over the display's control socket and get an ack. That is
the only way in: the cache-key mailboxes (`display_on_demand_request`,
`plugin_error_clear_request`) are gone, and a write to either is dropped
with a warning. When no display is listening yet, the start route starts the
service (if asked) and answers `202`; the web process's dispatcher
([`on_demand_dispatch.py`](../web_interface/on_demand_dispatch.py)) sends the
request once the socket is up, and the status routes report the outcome.
Any other failure is answered as an error. Socket commands and plugins'
in-process requests end in the same handler, `_handle_on_demand_request()`.
send the command over the display's control socket and get an ack; only when
the socket could not carry it (a stopped display, one older than the socket
or the command) do they write the mailbox instead. A display that had the
request and refused it is answered with the error, not posted a mailbox
copy. The display looks at the mailbox every
`MAILBOX_POLL_INTERVAL_WITH_SOCKET` (1 s) while it serves the socket, and
every `ON_DEMAND_POLL_INTERVAL` (0.25 s) without one, from its dwell sleep,
its render loops and Vegas's interrupt check as well as the main loop; a
look is one `stat()` unless the file changed. Both ways end in the same
handler, `_handle_on_demand_request()`.
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
for the protocol and the permission model.
for the protocol, the permission model and the plan to retire the mailboxes.
### Web and display processes: who runs plugins
@@ -117,8 +118,8 @@ other web-UI action runs its script as a subprocess. A later, explicit
The **control socket** from the web process to the display
([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) carries on-demand
commands, reloads an updated plugin and streams the display's state; it
replaced the cache-key mailboxes. The plugin web-entry
commands and reloads an updated plugin; its next stages stream the
display's state and retire the cache-key mailboxes. The plugin web-entry
contract above is still to come.
### Plugin state: desired, observed, and who owns it
+90 -119
View File
@@ -1,18 +1,18 @@
# Control socket (web → display)
The display process serves a Unix socket that the web interface uses to send
it commands and get an answer back. It replaced the cache-file "mailboxes" on
it commands and get an answer back. It replaces the cache-file "mailboxes" on
the SD card one command at a time. Stage 1 carries on-demand start, stop and
status. Stage 2 makes those commands land within a frame on every kind of
screen, and adds `brightness.set` and `plugin.reload`. Stage 3 adds a state
stream (`state.get`, `state.subscribe`), so the web interface reads what the
display is doing from the socket instead of from cache files the display
display wrote to the SD card. Stage 4 makes the socket the only way a command goes
while it works, and `errors.clear` replaces the last command that always
went through a mailbox. Stage 5 removes the mailboxes: the web interface no
longer writes them and the display no longer reads them, so the socket is
the only way a command reaches the display (see "Without the socket"). The
cache keys the display writes for readers stay as their fallback.
wrote to the SD card. Stage 4 makes the socket the only way a command goes
while it works: the web interface writes a mailbox only when the socket
cannot carry the request, `errors.clear` replaces the last command that
always went through a mailbox, and the display looks at the mailboxes once
a second, with a `stat()`. The file mailboxes and the cache keys stay as a
fallback for one release.
| | |
|---|---|
@@ -145,17 +145,14 @@ than the display's wait, so `pending` arrives before the client gives up.
`v` the display does not speak gets `unsupported_version`. `hello` is checked
by its `versions` list instead, and its result names the highest version both
sides share, so a client can find out what a display supports before it
relies on anything newer. The client sends `v: 1`, and a route answers an
error when the display refuses it. It does not send `hello` first, which
relies on anything newer. The client sends `v: 1` and falls back to the
mailbox when the display refuses it. It does not send `hello` first, which
saves a round trip.
New commands are added within a version, so stage 2 is still version 1. A
display that does not know a command answers `unknown_command`, which the
web interface treats like any other socket failure and falls back from, and
`hello` lists the commands a display knows.
(Since stage 5, a command the display does not know is an error the route
reports, telling the user to restart the display; there is no mailbox left
to fall back to.) The version changes only when the
`hello` lists the commands a display knows. The version changes only when the
envelope or the meaning of an existing command changes.
**Events.** `state.subscribe` is the one command with more than one message
@@ -220,8 +217,7 @@ socket.
}}
```
- `display` and `on_demand` are the dicts the cache keys hold (`on_demand`
includes `request_id`, the request it answers), `plugins` is
- `display` and `on_demand` are the dicts the cache keys hold, `plugins` is
the runtime snapshot (`build_runtime_snapshot`), and `brightness` is the
configured level, what the panel shows now, and whether the dim schedule
has it dimmed. A section not published yet is `null`.
@@ -372,12 +368,13 @@ request, validates it against the contract, and then does one of two things:
- For a query, it answers from a status snapshot the display provides
(`DisplayController._control_status`). The snapshot only reads attributes.
The render thread drains the queue in `_poll_on_demand_requests()`:
The render thread drains the queue in `_poll_on_demand_requests()`, the same
place it reads the mailbox:
- An on-demand command goes to `_handle_on_demand_request()`, which also
handles plugins' own requests. The two paths share all of their code:
activation, the request-id guard, error publishing, and resuming the
rotation afterwards.
- An on-demand command goes to `_handle_on_demand_request()`, which is the
mailbox's own handler. The two paths share all of their code: activation,
the processed-id guard, error publishing, and resuming the rotation
afterwards.
- `brightness.set` is applied there and then (`_apply_control_brightness`),
and the current frame is pushed again so the panel shows it.
- `plugin.reload` starts at the top of the next loop pass, the place where
@@ -410,13 +407,14 @@ The render thread drains the queue in `_poll_on_demand_requests()`:
unloads it mid-load; a disable saved meanwhile is applied once the
reload is done. A second reload of the same plugin runs after the first.
Draining the queue costs no disk read, so it has no floor. A queued command
also lets `_service_pending_changes()` skip its own 0.25 s floor.
The floor on the mailbox read (0.25 s, 1 s since stage 4) does not apply to the queue, because
draining it costs no disk read. A queued command also lets
`_service_pending_changes()` skip its own floor.
### Waking the render thread (stage 2)
Stage 1 made the socket answer, but not land sooner: a queued command waited
for the same polls the mailbox did. Measured on ledpi (Pi 4, 24 fps Vegas),
for the same polls the mailbox does. Measured on ledpi (Pi 4, 24 fps Vegas),
a start took 1.02 s on a static screen and about 0.4 s in Vegas either way.
Now the queue wakes the render thread:
@@ -436,7 +434,8 @@ Now the queue wakes the render thread:
- **Scrolling screens** already service pending changes every frame.
So a command lands within a millisecond or so on a static screen and in a
dwell, and within one frame in Vegas and on a scrolling screen. Commands still run only on the render thread: the
dwell, and within one frame in Vegas and on a scrolling screen. The mailbox
is slower on purpose (see "The mailboxes now"). Commands still run only on the render thread: the
connection threads only queue them and set the event. The one exception is
the slow half of `plugin.reload` (tearing down and loading the plugin),
which runs on its own thread. Every change to the display's state still
@@ -454,101 +453,74 @@ bookkeeping. A client's send to the render thread waking took 0.72 ms median
Without a socket (Windows, `LEDMATRIX_CONTROL_SOCKET=off`) the waits are the
plain sleeps they were.
**Exactly once.** A request goes over the socket once. The web route sends
it again only while no display is listening (see "Without the socket"), so
a display gets it at most once; a start with a request id the display has
just processed is dropped anyway (`on_demand_request_id`). The persisted
`display_on_demand_processed_id` guard against a mailbox replayed after a
restart went with the mailbox.
**Exactly once.** Since stage 4 the web interface writes the mailbox only
when the display never had the request (see "When the web interface falls
back"), so a request goes one way or the other, never both. A command and a
mailbox write for the same request still share one `request_id`, and the
`on_demand_request_id` and processed-id checks still drop a second copy: an
older web interface (before stage 4) wrote the mailbox after a reply timed
out, too. The display takes such a copy out of the mailbox when it drops it.
## Without the socket (stage 5)
## When the web interface falls back (stage 4)
The client tells a request the display never had from one it had and then
failed. `ControlError.sent` is True once the whole request was written to a
connected display; a refusal the display sends before reading anything
(`forbidden`, too many connections) carries no request id, and leaves it
False. `src.ipc.client.display_not_listening()` picks out the one case a
later retry can fix: nothing is listening (`no_socket`, `refused`) and the
request was never sent.
False. `src.ipc.client.should_fall_back()` is the one rule every route uses:
| What happened | Example reasons | On-demand start | On-demand stop | `errors.clear` |
|---|---|---|---|---|
| No display listening | `no_socket`, `refused` | service stopped: `400` without `start_service`; with it, start the service and answer `202` (`status: "starting"`) at once; the dispatcher sends the request until the display takes it (up to 45 s), else `start-timeout`. Service running (still starting): the same `202`, sent for up to 10 s | `503` ("not running" / "may still be starting"); with `stop_service` the service is stopped and the route succeeds | `503` ("not running"; its errors are the last run's, and the next run starts with none) |
| A display too old to know the command | `unknown_command`, `unsupported_version` | `503` | `503` | `503`, "restart it" |
| No socket in this process | `disabled`, `unsupported` (Windows, `LEDMATRIX_CONTROL_SOCKET=off`) | `503` | `503` | `503` |
| The display had it and failed, turned it away, or never answered | `busy`, `invalid_args`, `internal`, a timeout, a hang-up, `bad_response`, `forbidden` | `503` (`400` for `invalid_args`) | `503` (with `stop_service`: stopped anyway) | `503` |
| What happened | Example reasons | Mailbox? | The route answers |
|---|---|---|---|
| The display never had it | `no_socket`, `refused`, `disabled`, `unsupported`, a connect or send that timed out, `forbidden` / `busy` at the door, `invalid_request` (refused by the client itself) | yes | success, `transport: "mailbox"`, `socket_error` |
| A display too old to know it (the upgrade case) | `unknown_command`, `unsupported_version` | yes | as above |
| The display had it and failed | `busy` (queue full), `invalid_args`, `internal`, a timeout or hang-up after the send, `bad_response` | no | `503` (`400` for `invalid_args`), `socket_error` |
Every error answer carries `socket_error` (a reason code, or `other`).
Nothing is written to the cache in any of these cases.
**Waiting for a display that is starting.** The socket comes up when the
display's run loop starts, after every plugin has loaded, which can take
longer than a client waits (the MQTT bridge gives up after 15 s). So the
start route never waits: it answers `202` with `status: "starting"`, and
hands the request to the web process's one dispatcher
([`web_interface/on_demand_dispatch.py`](../web_interface/on_demand_dispatch.py)).
Its worker thread sends the request every 0.5 s while nothing is listening,
until the display acknowledges it or the wait runs out (45 s after a cold
start, `START_WAIT_SECONDS`; 10 s for a service that was already running,
`ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS`). Any other failure ends it at once.
One start is pending at a time: a newer start replaces it, and a stop
cancels it (the stop then succeeds even with no display listening, and
reports `cancelled_request_id`).
The outcome is reported where clients already look:
`GET /display/on-demand/status` answers the pending start's state
(`source: "web"`, `status: "starting"`, or `status: "error"` with `error:
"start-timeout"` or the socket's reason) until the display publishes
something newer, and `GET /display/current-status` adds it as
`on_demand_pending`.
A delivered start keeps reading as `status: "starting"`, now with
`delivered: true`, until the display publishes the state that answers it.
The display acknowledges a start as soon as its socket opens, but its run
loop acts on it only after the first screen is built (about 5 s on ledpi,
while Vegas renders its first strip), and meanwhile it publishes its own
idle state. The display's on-demand state names the request it answers
(`request_id`), so "answers it" means the id matches. A display older than
that field answers with any state published after the delivery. Either
way the delivered start is reported for at most 30 s
(`DELIVERED_SHOWN_SECONDS`).
A display that had the request may have applied it (a reply that timed out),
or would refuse the mailbox copy as well (bad arguments), or is stuck and
would not read the mailbox either (a full queue). Writing the copy anyway
only turned that into a "success". An on-demand stop with `stop_service`
still stops the service, which ends on-demand whatever happened.
Brightness and plugin reload never had a mailbox: without the socket, the
config watcher applies the saved brightness and a reload becomes the
restart banner, as before.
### The mailboxes are gone
### The mailboxes now
| Former mailbox | Last written by | Now |
|---|---|---|
| `display_on_demand_request` | the web interface on fallback (stage 4); plugins that predate `BasePlugin.request_on_demand()` | not read. A write is dropped by `CacheManager.save_cache` (`RETIRED_MAILBOX_KEYS`) and logged once per writer as a warning that names the plugin when it can be told (the plugin instance on the call stack, else the `plugin_id` in the request) |
| `plugin_error_clear_request` | the web interface on fallback (stage 4) | not read; a write is dropped and logged the same way |
| Mailbox | Written by | Read by the display | While the socket is up |
|---|---|---|---|
| `display_on_demand_request` | the web interface, only on fallback; plugins that predate `BasePlugin.request_on_demand()`, or run on a core without it | the render thread, `_poll_on_demand_requests()` | looked at every 1 s (`MAILBOX_POLL_INTERVAL_WITH_SOCKET`), 0.25 s without a socket |
| `plugin_error_clear_request` | the web interface, only on fallback | the error publisher's thread, every 5 s tick | unchanged rate |
A file left on the SD card by an older version is never read again, so it
is harmless. `MailboxWatch` and
`CacheManager.file_signature`, which made a look at a mailbox one `stat()`,
went with them.
A look is one `stat()` of the mailbox file (`CacheManager.file_signature`):
`(inode, mtime, size)`, and every write renames a new file into place, so a
new write always looks different. `MailboxWatch` reads the file only when
that changed since the last look, so a mailbox that holds nothing new, or
nothing at all, costs no open and no parse. A socket command never reads or
deletes the on-demand mailbox. A start already processed is taken out of
the mailbox instead of being re-read until it expires.
The four plugins that wrote `display_on_demand_request` (birdnet-go,
mqtt-notifications, on-air, pomodoro-timer) use `request_on_demand()` on a
core that has it, and fall back to the mailbox only when that method is
missing or answers `None` (no display in the process, or a full queue).
On a stage-5 core such a fallback write is dropped with the warning above.
A request that comes through the on-demand mailbox while the socket is up
is logged once per writer (`came through the file mailbox although the
control socket is up`), which names the plugins that still write it.
### Plugins in the display process
A plugin asks for the screen with `BasePlugin.request_on_demand()` and gives
it back with `end_on_demand()` (see "On-demand display" in
[PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md)). Neither goes through
the socket or a file: `PluginManager` hands the on-demand request,
the socket or a file: `PluginManager` hands the mailbox-shaped request,
marked `source: 'plugin'`, to `DisplayController.submit_plugin_on_demand`,
which queues it in memory (at most `PLUGIN_ON_DEMAND_QUEUE_SIZE`, 32) from
whatever thread the plugin called on, and wakes the render thread through
the socket's queue flag (`ControlServer.wake()`). The render thread applies
it in `_drain_control_commands`, after the socket's commands, through the
same `_handle_on_demand_request`, so it lands within a frame like a socket
command. Without a socket it lands on the next pending-changes pass. A
plugin's stop ends only a session that plugin owns.
command. Without a socket it lands on the next pending-changes pass (typically
within 0.25 s). A plugin's stop ends only a session that plugin owns. The four
plugins that wrote the mailbox (birdnet-go, mqtt-notifications, on-air,
pomodoro-timer) use it where the core has it and write the mailbox
otherwise.
## Robustness
@@ -566,7 +538,8 @@ block the render loop or crash it:
found. A client that disconnects mid-message is dropped silently. No
exception from a handler leaves the connection thread.
- **Full queue.** When the queue is full, the client gets `busy`, and the web
interface answers `503`. A full queue means the render thread
interface answers `503` rather than write the mailbox, which the stuck
render thread would not read either. A full queue means the render thread
is stuck, and the systemd watchdog deals with that.
- **Awaited commands.** The wait for an awaited command's outcome happens on
its connection thread and is bounded (`AWAIT_SECONDS`), so a stuck render
@@ -581,8 +554,8 @@ block the render loop or crash it:
process created.
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
as before. The web interface then cannot send it commands (the routes
answer `503`), and reads the cache keys and the heartbeat file.
as before. The web interface then uses the mailbox, and reads the cache
keys and the heartbeat file.
- **Subscribers (stage 3).** A `state.subscribe` connection gives its request
slot back and takes one of 4 subscriber slots (`MAX_SUBSCRIBERS`). A fifth
gets `busy`. So a few browsers' web processes holding streams can never
@@ -640,9 +613,11 @@ device never touches the live display.
## Stage plan
1. **On-demand, with acks (done, #706).** Contract, server, client.
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes tried
the socket first and reported `transport: "socket" | "mailbox"` (plus
`socket_error` on fallback).
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes try the
socket first and report `transport: "socket" | "mailbox"` (plus
`socket_error` on fallback). The mailbox is unchanged, and the plugins that
write it directly (birdnet-go, mqtt-notifications, on-air, pomodoro-timer)
keep working.
2. **Commands that were restarts or polls (done).**
- The render thread waits on the queue instead of sleeping, and Vegas
checks it every frame, so a command lands within a frame on every kind
@@ -685,12 +660,12 @@ device never touches the live display.
4. **The mailboxes become a fallback (done).**
- The web interface writes a mailbox only when the socket could not carry
the request (`should_fall_back`); a display that had it and failed is
answered as that.
answered as that (see "When the web interface falls back").
- `errors.clear` replaces `plugin_error_clear_request` as the way a clear
reaches the display.
- The display looks at the on-demand mailbox once a second while the
socket is up, reads either mailbox only when its file changed, and logs
who still writes the on-demand one.
who still writes the on-demand one (see "The mailboxes now").
- Not changed, deliberately: config saves (the schedule, the dim
schedule, plugin settings) still reach the display through
`config.json` and its watcher, which is the setting itself rather than
@@ -699,20 +674,15 @@ device never touches the live display.
already stats at most once a second. Plugin health and metrics resets
write the persisted record the display publishes and do not reach the
running display (their routes say so); they are not mailboxes.
5. **Remove the mailboxes (done, the release after 3.8.1).** The web
interface no longer writes `display_on_demand_request` or
`plugin_error_clear_request`, and the display no longer reads them (see
"Without the socket"). When no display is listening, the start route
starts the service if asked, answers `202`, and the web process's
dispatcher sends the request once the socket is up; every other failure
is an error the route reports. A write to
either key is dropped with a one-time warning naming the writer. The
display still writes `display_current_state`, `display_on_demand_state`
and `plugin_runtime_snapshot`: the web interface reads them whenever the
socket cannot answer (a stopped or starting display, a web user not yet
in the socket's group, a platform without Unix sockets), and
`display_on_demand_config` is the display's own record for resuming a
session after a restart. Retiring those keys is left for later.
5. **Remove the mailboxes (next release).** Once every device has run a
display with stage 4, the web interface stops writing both mailboxes and
the display stops reading them. The four plugins that wrote
`display_on_demand_request` now have an in-process way to ask for the
screen (`BasePlugin.request_on_demand()` / `end_on_demand()`, see
"Plugins in the display process"); they keep the mailbox write only as
their fallback on older cores. The display also stops writing `display_current_state`,
`display_on_demand_state` and `plugin_runtime_snapshot` once the web
interface no longer falls back to them.
## Checking it on a device
@@ -724,11 +694,12 @@ curl -s -X POST localhost:5000/api/v3/display/on-demand/start \
# ... "transport": "socket"
```
On an error, `socket_error` gives the reason. `no_socket` means the display
is stopped or still starting. `refused` usually means the web user is not in
the socket's group, which takes effect when the web service restarts after
the user is added. `busy`, `timeout` and the like mean the display had the
request and did not take it. Nothing is ever written to a mailbox.
If the response says `"transport": "mailbox"`, `socket_error` gives the
reason. `no_socket` means the display is stopped or predates the socket.
`refused` usually means the web user is not in the socket's group, which
takes effect when the web service restarts after the user is added. A `503`
with `"transport": "socket"` means the display had the request and did not
take it (`busy`, `timeout`, ...): nothing was written to the mailbox.
An error clear:
@@ -736,7 +707,7 @@ An error clear:
curl -s -X POST localhost:5000/api/v3/errors/clear \
-H 'Content-Type: application/json' -d '{"all":true}'
# ... "applied": true, "transport": "socket"
sudo journalctl -u ledmatrix | grep -E "Cleared .* plugin error|retired"
sudo journalctl -u ledmatrix | grep -E "Cleared .* plugin error|file mailbox"
```
Brightness and a plugin reload:
+20 -30
View File
@@ -522,52 +522,42 @@ Returns the request id once queued, or `None` as above.
These methods are new after core 3.8.0 (see `CHANGELOG.md`). Before them,
plugins wrote the `display_on_demand_request` cache key (the "mailbox")
themselves. **The mailbox is gone** (see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md), stage 5): the display no
longer reads it, and a write to it is dropped with a warning in the log,
once per plugin:
```
Ignored a write to the retired 'display_on_demand_request' cache key by plugin 'my-plugin': ...
```
A plugin that only needs to run on cores with these methods calls them and
treats `None` as "no display took it":
```python
def _show_alert(self):
if self.request_on_demand(mode="my_alert", duration=15) is None:
self.logger.info("No display to show the alert on")
```
A plugin that must also work on cores before them can keep the mailbox
write as its fallback, guarded by `hasattr`: on those cores the display
still reads it, and on a current core the write is only dropped and logged
(when the method is missing it never runs at all). Raise
`ledmatrix_min_version` to the release that added the methods once you no
longer need it.
themselves. The display reads it only once a second while the control
socket is up, and it will be removed in a future release (see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md), stage 5). A plugin that
must keep working on older cores checks for the method, and writes the
mailbox only when the method is missing or answers `None`:
```python
import time, uuid
def _show_alert(self):
if hasattr(self, "request_on_demand"):
self.request_on_demand(mode="my_alert", duration=15)
if hasattr(self, "request_on_demand") and self.request_on_demand(
mode="my_alert", duration=15):
return
# A core older than request_on_demand(): the mailbox it still reads.
# Older core, or no display in this process: the mailbox, as before.
self.cache_manager.set("display_on_demand_request", {
"request_id": str(uuid.uuid4()), "action": "start",
"plugin_id": self.plugin_id, "mode": "my_alert",
"duration": 15, "pinned": False, "timestamp": time.time(),
})
def _release(self):
if hasattr(self, "end_on_demand") and self.end_on_demand():
return
self.cache_manager.set("display_on_demand_request", {
"request_id": str(uuid.uuid4()), "action": "stop",
"plugin_id": self.plugin_id, "timestamp": time.time(),
})
```
`end_on_demand()` ends only the plugin's own session; a stop written to the
mailbox on an older core ends any session, whoever started it.
Keep `ledmatrix_min_version` where it is: the fallback is what keeps the
plugin working on older cores. A mailbox stop ends any on-demand session,
whoever started it; `end_on_demand()` ends only the plugin's own.
Both methods answer a request id only when the plugin manager returned a
string, so a test that gives the plugin a `MagicMock()` plugin manager gets
`None`. To test the path where the display takes the request, set
`None` and exercises the mailbox path. To test the new path, set
`plugin_manager.request_on_demand.return_value = "some-id"`.
> The full source for `BasePlugin` lives in
+34 -60
View File
@@ -464,7 +464,7 @@ Request a specific plugin to display on-demand.
- `mode` (string, optional): Display mode name (plugin_id inferred if not provided)
- `duration` (number, optional): Duration in seconds (0 = until stopped)
- `pinned` (boolean, optional): Pin display (pause rotation)
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket. A stopped one is started, and the route answers `202` at once (see below); the request is sent once the display's socket is up, which can take as long as the display takes to load its plugins. When false and the service is stopped, the route returns 400.
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket (within about a second through the mailbox fallback). When false and the service is stopped, the route returns 400.
**Response**:
```json
@@ -484,54 +484,24 @@ Request a specific plugin to display on-demand.
`service` is `null` when `start_service` is false.
`transport` is always `"socket"`: the display's control socket acknowledged
the request (it is queued for the render thread, which wakes for it and
applies it within a frame; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)).
The `"mailbox"` value earlier releases could answer is gone with the
mailbox: nothing is written to the cache.
`transport` says how the request reached the display: `"socket"` means the
display's control socket acknowledged it (it is queued for the render thread,
which wakes for it and applies it within a frame; see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), `"mailbox"` means it was
written to the cache mailbox the display polls, as before the socket existed.
The mailbox is used only when the socket could not carry the request. With
`"mailbox"`, `socket_error` gives the reason (`no_socket` when the display is
stopped or predates the socket, `refused`, a connect `timeout`,
`unknown_command` from a display too old for the command, ...). Either way
the request is applied the same way; `request_id` is the same id in both.
**No display listening yet** (`no_socket`, `refused`: the service is
stopped, or still loading its plugins). With the service stopped and
`start_service` false, `400`. Otherwise the route starts the service if
needed and answers at once:
```json
{
"status": "starting",
"message": "The display service is starting; ...",
"data": {"request_id": "uuid-here", "plugin_id": "football-scoreboard", "mode": "nfl_live",
"duration": 45, "pinned": true, "service": {"active": true, "started": true},
"transport": "socket", "socket_error": "no_socket",
"pending": true, "wait_seconds": 45.0}
}
```
with HTTP `202`. The web process sends the request until the display takes
it, for up to `wait_seconds` (45 after a cold start, 10 when the service was
already running). Follow it with `GET /api/v3/display/on-demand/status`:
its `state` is `{status: "starting", source: "web", request_id, ...}` while
it waits, and still `starting` with `delivered: true` once the display has
acknowledged it but not yet published the state for that `request_id` (at
most 30 s), then the display's own state, or `{status: "error",
error: "start-timeout"}` (or the socket's reason) if it never was;
`GET /api/v3/display/current-status` carries the same as
`on_demand_pending`. A newer start replaces a pending one; a stop cancels it.
Otherwise, when the display did not take the request, the route answers an
error with `status: "error"` and `data: {request_id, transport: "socket",
socket_error}`:
- a full queue (`busy`), no answer after the request was sent (`timeout`,
`closed`), a display older than the command (`unknown_command`), no
socket in the web process (`disabled`, `unsupported`): `503` at once;
- bad arguments (`invalid_args`): `400`.
The stop route answers the same errors (`503` when no display is
listening, with a message saying whether the service is stopped), except
that with `stop_service: true` it still stops the service and answers
success, with `socket_error` set, and that a stop which cancelled a pending
start succeeds with `cancelled_request_id` even when no display is
listening.
When the display had the request and did not take it -- a full queue
(`busy`), bad arguments (`invalid_args`), no answer after the request was
sent (`timeout`, `closed`) -- the route answers `503` (`400` for
`invalid_args`) with `status: "error"` and `data: {request_id, transport:
"socket", socket_error}`, and writes nothing to the mailbox. The stop route
does the same, except that with `stop_service: true` it still stops the
service and answers success.
### Stop On-Demand Display
@@ -561,7 +531,7 @@ Stop the current on-demand display.
}
```
`transport` and the errors are as for start; a stop is never retried.
`transport` and `socket_error` are as for start.
---
@@ -2227,7 +2197,7 @@ Every response below adds three fields to the shape it always had:
|---|---|
| `snapshot_available` | `false` until the display service has reported (for example, it is not running). Counts are then zero. |
| `generated_at` | When the display service produced the snapshot (ISO, the Pi's local time), or `null`. |
| `clear_pending` | Always `false`: a clear is applied before its route answers. Kept for compatibility. |
| `clear_pending` | A clear has been requested and the display service has not applied it yet. |
### Get Error Summary
@@ -2308,17 +2278,21 @@ before it answers: `applied` is `true`, `transport` is `"socket"`, and
}
```
`clear_requested`, `applied` and `transport` are always `true`, `true` and
`"socket"`, and are kept for compatibility.
When the socket cannot carry it (the display is stopped, or older than
`errors.clear`) the clear is asynchronous, as before: the web interface
records a request (`plugin_error_clear_request` in the shared cache),
`applied` is `false` and `transport` is `"mailbox"`, and the display service
applies it within about 5 seconds. Reads hide the cleared errors from the
moment the request is recorded. Until the display service applies an
age-based clear, `recent_errors` and `active_patterns` are already filtered
but the counts are the old ones, and `clear_pending` is `true`. Then
`cleared_count` is how many of the reported errors the clear hides, and
`null` when that cannot be known before the display service applies it (an
age-based clear over more errors than the report lists).
When the socket cannot carry the clear, the route answers `503` with
`context.socket_error`, and nothing is cleared or recorded: the
`plugin_error_clear_request` mailbox earlier releases fell back to is gone.
The message says why: the display service is not running (its errors are
then the last run's, and its next run starts with none), it is too old for
`errors.clear` (restart it), the web process has no socket, or the display
had the request and failed it (`internal`, a timeout after the request was
sent).
A request that could not be written to the shared cache answers `500`. A
display that had the request and failed it (`internal`, a timeout after the
request was sent) answers `503`, with `context.socket_error`.
---
+48 -54
View File
@@ -25,11 +25,10 @@ Typical plugin usage::
import json
import os
import sys
import time
from datetime import datetime
import pytz
from typing import Any, Dict, List, Optional
from typing import Any, Dict, List, Optional, Tuple
import logging
import threading
import tempfile
@@ -73,60 +72,41 @@ def _outlived(record: Any, max_age: Optional[float], now: float) -> bool:
return False
#: Cache keys that were file "mailboxes" from the web interface (and some
#: plugins) to the display. The control socket replaced them, and nothing
#: reads them any more, so a write is refused rather than left on the SD card
#: for nobody: see :func:`_refuse_retired_mailbox_write`.
RETIRED_MAILBOX_KEYS = frozenset({'display_on_demand_request', 'plugin_error_clear_request'})
#: (key, writer) pairs already warned about, so a plugin that writes on every
#: event logs once per process, not once per write.
_retired_writers_warned: set = set()
_retired_writers_lock = threading.Lock()
_NOT_SEEN: Any = object()
def _retired_mailbox_writer(data: Any) -> str:
"""Name whoever is writing a retired mailbox key, as well as can be told.
class MailboxWatch:
"""Tells the poller of a mailbox key whether its file changed since the
last look, from one stat() (:meth:`CacheManager.file_signature`).
The plugin instance on the call stack when there is one (a ``self`` with
a string ``plugin_id`` and a ``cache_manager``: what BasePlugin gives
every plugin), else the ``plugin_id`` the request itself names, else
``'unknown'``.
The display polls the mailboxes the web interface falls back to. Reading
one is an open and a JSON parse; with this a poll that finds the same file
(or none) costs a stat, and the file is read only after a new write. A
cache without ``file_signature`` (a test double) is read every time.
"""
frame = sys._getframe(2) # pylint: disable=protected-access
depth = 0
while frame is not None and depth < 25:
owner = frame.f_locals.get('self')
plugin_id = getattr(owner, 'plugin_id', None) if owner is not None else None
if isinstance(plugin_id, str) and plugin_id and hasattr(owner, 'cache_manager'):
return f"plugin '{plugin_id}'"
frame = frame.f_back
depth += 1
payload = data.get('data', data) if isinstance(data, dict) else None
named = payload.get('plugin_id') if isinstance(payload, dict) else None
if isinstance(named, str) and named:
return f"plugin '{named}' (named in the request)"
return 'unknown'
def __init__(self, key: str):
self.key = key
self._seen: Any = _NOT_SEEN
def _refuse_retired_mailbox_write(logger: logging.Logger, key: str, data: Any) -> None:
"""Warn, once per writer, that a write to a retired mailbox key was dropped."""
try:
writer = _retired_mailbox_writer(data)
except Exception: # pylint: disable=broad-except
writer = 'unknown'
with _retired_writers_lock:
if (key, writer) in _retired_writers_warned:
return
_retired_writers_warned.add((key, writer))
if key == 'display_on_demand_request':
hint = ("call self.request_on_demand() / self.end_on_demand() instead "
"(BasePlugin, LEDMatrix 3.8.1 and later)")
else:
hint = "clear errors through POST /api/v3/errors/clear instead"
logger.warning("Ignored a write to the retired '%s' cache key by %s: the display no "
"longer reads this file mailbox. Update it to %s. (Logged once per writer.)",
key, writer, hint)
def changed(self, cache_manager: Any) -> bool:
"""True when the poller should read the key now."""
signature = getattr(cache_manager, 'file_signature', None)
sig = signature(self.key) if callable(signature) else _NOT_SEEN
if sig is not None and not isinstance(sig, tuple):
return True # cannot tell: read it
if sig is None:
self._seen = None
return False # no file, nothing to read
if sig == self._seen:
return False
self._seen = sig
return True
def forget(self) -> None:
"""Read the key on the next poll even if its file has not changed
(the last read failed)."""
self._seen = _NOT_SEEN
class CacheManager:
@@ -353,6 +333,24 @@ class CacheManager:
"""Get the path for a cache file."""
return self._disk_cache_component.get_cache_path(key)
def file_signature(self, key: str) -> Optional[Tuple[int, int, int]]:
"""``(st_ino, st_mtime_ns, st_size)`` of ``key``'s file, or None when
there is no file (the key is absent, or this cache has no disk tier).
One stat(), no read: a poller of a mailbox another process writes
compares it with the last one it saw and reads the file only when it
changed. Every write replaces the file (a temp file renamed into
place), so a new write always has a new inode, however fast it came.
"""
path = self._get_cache_path(key)
if not path:
return None
try:
st = os.stat(path)
except OSError:
return None
return (st.st_ino, st.st_mtime_ns, st.st_size)
def get_cached_data(self, key: str, max_age: int = 300, memory_ttl: Optional[int] = None) -> Optional[Dict[str, Any]]:
"""Get data from cache (memory first, then disk) honoring TTLs.
@@ -390,10 +388,6 @@ class CacheManager:
key: Cache key
data: Data to cache
"""
if key in RETIRED_MAILBOX_KEYS:
_refuse_retired_mailbox_write(self.logger, key, data)
return
# Periodic cleanup before adding new entries
self._cleanup_memory_cache()
+183 -36
View File
@@ -49,7 +49,7 @@ from src.screen_runner import (
from src.display_manager import DisplayManager
from src.config_manager import ConfigManager
from src.config_service import ConfigService
from src.cache_manager import CacheManager
from src.cache_manager import CacheManager, MailboxWatch
from src.font_manager import FontManager
from src.logging_config import get_logger
from src.exceptions import PluginError
@@ -69,6 +69,10 @@ from src.vegas_mode.render_pipeline import SYNC_SEND_INTERVAL
# Get logger with consistent configuration
logger = get_logger(__name__)
# The on-demand file mailbox: the fallback for a web interface that cannot
# reach the control socket, and how some plugins still ask for the screen.
ON_DEMAND_MAILBOX_KEY = 'display_on_demand_request'
# How often the unchanged current mode is republished for the web UI, which
# treats display_current_state older than 120 s as unknown.
CURRENT_STATE_REFRESH_SECONDS = 30
@@ -421,15 +425,19 @@ class DisplayController:
# coordinator exists (Vegas was off at startup). The render thread
# creates it in _is_vegas_mode_active(), never the watcher thread.
self._pending_vegas_init = False
# Monotonic stamp of the last _service_pending_changes pass. None
# means "never", so the first call always goes through.
# Monotonic stamp of the last mailbox disk read; see
# _poll_on_demand_requests. None means "never polled", so the first
# call always goes through.
self._last_on_demand_poll: Optional[float] = None
# Monotonic stamp of the last _service_pending_changes pass; same
# "None means never" convention as _last_on_demand_poll.
self._last_pending_service: Optional[float] = None
# Monotonic stamp of the last scheduled-update pass; see
# _tick_plugin_updates_if_due. Same "None means never" convention.
self._last_plugin_update_tick: Optional[float] = None
# The control socket (src/ipc), started by run(). None when it is not
# served (Windows, LEDMATRIX_CONTROL_SOCKET=off, a bind failure);
# then only plugins in this process can start on-demand sessions.
# the file mailbox works either way.
self._control_server = None
# A brightness set_brightness() refused, so the periodic service pass
# doesn't retry (and log) the same failure several times a second.
@@ -1791,9 +1799,6 @@ class DisplayController:
'status': self.on_demand_status,
'error': self.on_demand_last_error,
'last_event': self.on_demand_last_event,
# The request this state answers: lets the web interface tell
# the outcome of a start it delivered from an older state.
'request_id': self.on_demand_request_id,
'remaining': self._get_on_demand_remaining(),
'last_updated': time.time()
}
@@ -1862,16 +1867,34 @@ class DisplayController:
self.force_change = True
self._publish_on_demand_state()
#: Shortest gap between _service_pending_changes passes. The schedule
#: checks are gated to once per clock minute and the rest is attribute
#: compares. Callers run at frame rate,
#: Shortest gap between mailbox disk reads. This is called after every
#: frame -- about 125 times a second on a scrolling mode -- and the read
#: below is deliberately uncached, so without a floor it was 125 disk reads
#: per second to find nothing. An on-demand request comes from a person
#: clicking in the web UI, so a quarter second of latency is not
#: perceptible, and it cuts the read rate by 30x.
ON_DEMAND_POLL_INTERVAL = 0.25
#: The mailbox poll while the control socket is up. The web interface then
#: writes the mailbox only when it could not reach the socket (a display
#: being restarted, a web user not yet in the socket's group), and the
#: plugins that still write it get the screen within this long. A look is
#: one stat() of the mailbox file (MailboxWatch).
MAILBOX_POLL_INTERVAL_WITH_SOCKET = 1.0
#: Shortest gap between _service_pending_changes passes. The same floor
#: the mailbox poll had before the socket, since that read was the only
#: real cost in the pass: the schedule checks are gated to once per clock
#: minute and the rest is attribute compares. Callers run at frame rate,
#: so between passes the whole cost is one monotonic-clock compare.
PENDING_CHANGES_INTERVAL = 0.25
#: Class-level defaults for controllers built without __init__ (tests).
_control_server: Optional[ControlServer] = None
#: The last on-demand request handled; published in _on_demand_state.
on_demand_request_id: Optional[str] = None
#: Created on the first poll; see _poll_on_demand_requests.
_on_demand_mailbox: Optional[MailboxWatch] = None
#: Writers whose mailbox requests have been logged (_note_mailbox_request).
_mailbox_writers_logged: FrozenSet[str] = frozenset()
#: Most plugin on-demand requests waiting for the render thread at once.
#: A plugin that asks faster than the display drains (four times a
#: second at worst) is refused, not queued without end.
@@ -2014,11 +2037,44 @@ class DisplayController:
on_demand_plugin_id, len(enabled_plugins))
return enabled_plugins
def _consume_on_demand_request(self, request_id: str) -> None:
"""Remove the request we just handled from the mailbox.
Leaving it on disk meant a restart replayed the previous request: the
fresh controller read it, activated it and cached it, so the request
the caller had just made was ignored and the panel silently showed the
earlier plugin.
Compare before deleting. The web process can post a newer request
between the read and this delete; an unconditional delete threw that
one away and it was never processed -- the user's second click did
nothing. Re-reading uncached and only deleting our own request_id
leaves a newer request in the mailbox for the next poll instead.
This narrows the window rather than closing it: a request landing
between the re-read and the delete is still lost. Closing it properly
needs an atomic claim (a rename, or a compare-and-delete primitive)
that the cache layer does not currently offer, so the honest fix is a
smaller window plus this note, not a bigger lock. For start requests
processed_id still guards against reprocessing if the delete fails.
"""
try:
current = self.cache_manager.get(ON_DEMAND_MAILBOX_KEY,
max_age=3600, memory_ttl=0)
if not current or current.get('request_id') == request_id:
self.cache_manager.delete(ON_DEMAND_MAILBOX_KEY)
else:
logger.debug("Newer on-demand request %s arrived while processing "
"%s; leaving it in the mailbox",
current.get('request_id'), request_id)
except (OSError, AttributeError, KeyError) as err:
logger.debug("Could not clear the on-demand request mailbox: %s", err)
def _start_control_server(self) -> None:
"""Serve the control socket (src/ipc/server.py). Never raises.
Its handlers only queue commands; _drain_control_commands applies
them on the render thread.
them on the render thread, where the mailbox is read.
"""
if self._control_server is not None:
return
@@ -2031,8 +2087,7 @@ class DisplayController:
state_hub=hub,
handlers={ControlCommand.ERRORS_CLEAR: apply_error_clear})
except Exception: # pylint: disable=broad-except
logger.exception("Control socket not started; the web interface cannot "
"send this display commands")
logger.exception("Control socket not started; using the file mailbox only")
if self._control_server is not None:
self._start_state_stream(hub)
@@ -2060,8 +2115,10 @@ class DisplayController:
def _drain_control_commands(self) -> None:
"""Apply the commands that arrived over the control socket.
On-demand commands go through _handle_on_demand_request, as plugins'
own requests do, so both ways in behave the same. A brightness is applied
On-demand commands go through _handle_on_demand_request, the
mailbox's own handler, so both ways in behave the same, and a
request that came both ways (a client that timed out and fell back)
has one request id and is processed once. A brightness is applied
here. A plugin reload waits for the top of the next loop pass, where
no plugin is on the stack (_apply_pending_plugin_reloads); until
then the current screen ends early (_plugin_reload_pending).
@@ -2094,7 +2151,7 @@ class DisplayController:
``PluginManager.request_on_demand`` / ``end_on_demand`` (which
BasePlugin's methods of the same names call) build ``request``: the
on-demand request shape, with ``source: 'plugin'`` and the asking plugin's
mailbox's shape, with ``source: 'plugin'`` and the asking plugin's
id. Nothing here touches the panel or the on-demand state; the render
thread applies the request where it applies a socket command
(_drain_control_commands), through _handle_on_demand_request, and is
@@ -2389,35 +2446,105 @@ class DisplayController:
}
command.succeed(dict(result))
def _poll_on_demand_requests(self) -> None:
"""Apply on-demand requests: the control socket's, then plugins' own.
def _mailbox_poll_interval(self) -> float:
"""How often the on-demand mailbox is looked at: its old 0.25 s when
it is the only way in, MAILBOX_POLL_INTERVAL_WITH_SOCKET while the
control socket carries the web interface's commands."""
if self._control_server is not None:
return self.MAILBOX_POLL_INTERVAL_WITH_SOCKET
return self.ON_DEMAND_POLL_INTERVAL
Both are in memory (no disk read), so this has no floor and is
cheap to call every frame. The file mailbox
(``display_on_demand_request``) is gone: nothing reads it, and a
write to it is dropped with a warning (CacheManager).
def _poll_on_demand_requests(self) -> None:
"""Apply on-demand requests: the control socket's, then the mailbox's.
Socket commands are in memory and are applied at once. The file
mailbox (``display_on_demand_request``) is the fallback for a web
interface that could not reach the socket, and the way four plugins
still ask for the screen. It is looked at once per poll interval
(_mailbox_poll_interval), and read only when its file changed
(MailboxWatch): a look that finds nothing new is one stat().
"""
# Socket commands are already in memory: no disk read, so no floor.
self._drain_control_commands()
now = time.monotonic()
if (self._last_on_demand_poll is not None
and now - self._last_on_demand_poll < self._mailbox_poll_interval()):
return
self._last_on_demand_poll = now
watch = self._on_demand_mailbox
if watch is None:
watch = self._on_demand_mailbox = MailboxWatch(ON_DEMAND_MAILBOX_KEY)
if not watch.changed(self.cache_manager):
return
try:
# Use a long max_age (1 hour) to ensure requests aren't expired before processing
# The request_id check prevents duplicate processing.
#
# memory_ttl=0 is required, not optional: this key is a mailbox the
# web process writes and this process reads. get() defaults the
# in-memory TTL to max_age, so without it the first request read was
# pinned in memory for the full hour and every later poll returned
# that stale copy -- meaning no second on-demand request was honoured
# for an hour, while the API still reported success.
request = self.cache_manager.get(ON_DEMAND_MAILBOX_KEY,
max_age=3600, memory_ttl=0)
except (OSError, RuntimeError, ValueError, TypeError) as err:
watch.forget() # read it again next time
logger.error("Failed to read on-demand request: %s", err, exc_info=True)
return
if not isinstance(request, dict):
return
self._note_mailbox_request(request)
self._handle_on_demand_request(request)
def _note_mailbox_request(self, request: Dict[str, Any]) -> None:
"""Log, once per writer, an on-demand request that came through the
mailbox while the control socket is up.
The web interface writes the mailbox only when the socket fails, so
this is mostly a plugin that writes ``display_on_demand_request``
itself. The mailbox is going away; the log says who still uses it.
"""
if self._control_server is None:
return
writer = request.get('plugin_id') or request.get('mode') or 'unknown'
if not isinstance(writer, str):
writer = 'unknown'
if writer in self._mailbox_writers_logged:
return
self._mailbox_writers_logged = self._mailbox_writers_logged | {writer}
logger.info("On-demand %s request %s (for %s) came through the file mailbox "
"although the control socket is up. The mailbox is deprecated: it "
"is read every %.1fs and will be removed in a future release.",
request.get('action'), request.get('request_id'), writer,
self.MAILBOX_POLL_INTERVAL_WITH_SOCKET)
def _handle_on_demand_request(self, request: Dict[str, Any]) -> None:
"""Process one on-demand request, from the control socket or a plugin.
"""Process one on-demand request, from the mailbox or the control socket.
A socket command carries ``source: 'socket'``, and a plugin's own
request (submit_plugin_on_demand) ``source: 'plugin'``.
request (submit_plugin_on_demand) ``source: 'plugin'``. Only a
mailbox request is removed from the mailbox afterwards: the others
never put anything there, so that would be a disk read and maybe a
delete for nothing.
A plugin's stop ends only that plugin's own session: a plugin
releasing the screen must not end one the user started for
another plugin. A stop from the socket (the web interface) ends any
session.
another plugin. (A stop through the mailbox ends any session, as it
always has.)
"""
request_id = request.get('request_id')
if not request_id:
return
source = request.get('source')
from_mailbox = source not in ('socket', 'plugin')
action = request.get('action')
# For stop requests, always process them (no request-id guard)
# For stop requests, always process them (don't check processed_id)
# This allows stopping even if the same stop request was sent before
if action == 'stop':
if source == 'plugin' and not (
@@ -2443,22 +2570,42 @@ class DisplayController:
# without this the status route kept reporting it until
# the state aged out (120s) or another request came in.
self._clear_on_demand(reason='requested-stop')
# Stop requests are deliberately exempt from the request_id
# guard below, so that a second click stops a mode that a race
# left running.
# Stop requests are deliberately exempt from the request_id/
# processed_id guards above, so that a second click stops a mode
# that a race left running. Consuming the mailbox is therefore the
# only thing that ends the request: without it the same stop was
# re-read and re-processed on every poll, forever, logging at
# ON_DEMAND_POLL_INTERVAL for the life of the process.
if from_mailbox:
self._consume_on_demand_request(request_id)
return
# A start already processed (a client that sent the same request
# id twice) is not applied a second time.
# For start requests, check if already processed. A duplicate in the
# mailbox (a copy of a socket command, or one read before a restart)
# is taken out of it too, so it is not read again.
if request_id == self.on_demand_request_id:
logger.debug("On-demand start request %s already processed", request_id)
logger.debug("On-demand start request %s already processed (instance check)", request_id)
if from_mailbox:
self._consume_on_demand_request(request_id)
return
# Also check persistent processed_id (for restart scenarios)
processed_request_id = self.cache_manager.get('display_on_demand_processed_id', max_age=3600)
if request_id == processed_request_id:
logger.debug("On-demand start request %s already processed (persisted check)", request_id)
if from_mailbox:
self._consume_on_demand_request(request_id)
return
logger.info("Received on-demand request %s: %s (plugin_id=%s, mode=%s, via %s)",
request_id, action, request.get('plugin_id'), request.get('mode'),
source or 'unknown')
'mailbox' if from_mailbox else source)
# Mark as processed BEFORE processing (to prevent duplicate processing)
self.cache_manager.set('display_on_demand_processed_id', request_id, ttl=3600)
self.on_demand_request_id = request_id
if from_mailbox:
self._consume_on_demand_request(request_id)
if action == 'start':
logger.info("Processing on-demand start request for plugin: %s", request.get('plugin_id'))
+175 -54
View File
@@ -19,7 +19,7 @@ import uuid
from collections import defaultdict
from dataclasses import dataclass, field
from datetime import datetime, timedelta
from typing import Dict, List, Optional, Any, Callable
from typing import Dict, List, Optional, Any, Callable, Tuple
import logging
from src.exceptions import LEDMatrixError
@@ -487,28 +487,33 @@ def record_error(
# service publishes to the shared cache directory -- the same channel, and the
# same file permissions, as display_current_state and plugin_metrics_snapshot: files
# are 0660 and carry the cache directory's group, so root writes and the web
# user reads.
# user reads, and the other way round for the clear request.
#
# ERROR_SNAPSHOT_KEY written by the display service only
# ERROR_CLEAR_REQUEST_KEY written by the web interface only, as a fallback
#
# A clear goes over the control socket (``errors.clear``): the display applies
# it (clear_before) and republishes the snapshot before it answers. When the
# socket cannot carry it, the clear fails and the route says so: the
# ``plugin_error_clear_request`` file mailbox it used to fall back to is gone.
# A stopped display's errors go anyway: its next run publishes an empty
# snapshot over the old one. The web interface never writes the snapshot
# itself: two writers would race, and a snapshot owned by the web user is one
# more file root's write has to replace.
# it (clear_before) and republishes the snapshot before it answers. Only when
# the socket cannot carry it (no socket, or a display older than the command)
# does the web interface record a request in the mailbox, which the display
# applies on its next tick; its tick reads that file only when it changed.
# Until it has, the web interface hides whatever the snapshot shows from
# before the cutoff, so a clear takes effect for readers immediately and a
# snapshot published just before the request cannot bring old errors back.
# The web interface never writes the snapshot itself: two writers would race,
# and a snapshot owned by the web user is one more file root's write has to
# replace.
ERROR_SNAPSHOT_KEY = "plugin_error_snapshot"
ERROR_CLEAR_REQUEST_KEY = "plugin_error_clear_request"
#: Shortest gap between two snapshot writes, in seconds. A plugin failing in
#: a tight loop changes the aggregator many times a second; the snapshot is
#: rewritten at most this often, and only when something changed.
SNAPSHOT_MIN_INTERVAL = 10.0
#: How often the display service checks for changes. A check is an
#: in-memory comparison.
#: How often the display service checks for changes and clear requests. A
#: check is an in-memory comparison plus reading one small file.
SNAPSHOT_TICK_INTERVAL = 5.0
_SNAPSHOT_RECENT_ERRORS = 20
@@ -577,8 +582,8 @@ class ErrorSnapshotPublisher:
Runs in the display service only. tick() is the whole job; start() just
calls it from a daemon thread every SNAPSHOT_TICK_INTERVAL seconds, which
also means errors recorded while a write was being throttled still reach
the cache once the interval has passed. A clear (``errors.clear`` over
the control socket) is applied by :meth:`clear_now`.
the cache once the interval has passed, and a clear request is applied
even when no new error arrives to trigger a publish.
Nothing here raises: a failure to read or write the cache is logged at
debug and retried on a later tick.
@@ -596,6 +601,11 @@ class ErrorSnapshotPublisher:
self._published_version: Optional[int] = None
self._last_attempt: Optional[float] = None
self._applied_clear_id: Optional[str] = None
# The widest cutoff applied in this process: a clear request at or
# before it has nothing left to clear (see _pending_cutoff).
self._applied_clear_cutoff: Optional[float] = None
from src.cache_manager import MailboxWatch # the display's cache, loaded already
self._mailbox = MailboxWatch(ERROR_CLEAR_REQUEST_KEY)
self._tick_lock = threading.Lock()
self._stop = threading.Event()
self._thread: Optional[threading.Thread] = None
@@ -607,9 +617,39 @@ class ErrorSnapshotPublisher:
cleared = self.aggregator.clear_before(datetime.fromtimestamp(cutoff))
_snapshot_logger.info("Cleared %d plugin error record(s) as requested (%s)",
cleared, request_id)
if self._applied_clear_cutoff is None or cutoff > self._applied_clear_cutoff:
self._applied_clear_cutoff = cutoff
# A malformed request is acknowledged too, so it is not retried forever.
self._applied_clear_id = request_id
return cleared
def _apply_clear_request(self) -> bool:
"""Honour a mailbox clear request we have not applied yet. True if one was.
The mailbox is the fallback for a web interface that could not use
the control socket (``errors.clear``, :meth:`clear_now`). It is read
only when its file changed since the last tick; otherwise a tick
costs one stat().
"""
if not self._mailbox.changed(self.cache_manager):
return False
try:
request = self.cache_manager.get(ERROR_CLEAR_REQUEST_KEY, max_age=None, memory_ttl=0)
except Exception:
self._mailbox.forget()
raise
if not isinstance(request, dict):
return False
request_id = request.get("request_id")
if not isinstance(request_id, str) or not request_id or request_id == self._applied_clear_id:
return False
try:
cutoff = float(request.get("cutoff"))
except (TypeError, ValueError):
cutoff = float("nan")
self._clear(request_id, cutoff)
return True
def clear_now(self, request_id: str, cutoff: float) -> int:
"""``errors.clear`` over the control socket: apply a clear at once and
republish the snapshot, so the web interface's next read has it.
@@ -627,20 +667,23 @@ class ErrorSnapshotPublisher:
self._last_attempt = now
snapshot = self.aggregator.build_snapshot()
snapshot["applied_clear_id"] = self._applied_clear_id
snapshot["applied_clear_cutoff"] = self._applied_clear_cutoff
self.cache_manager.set(ERROR_SNAPSHOT_KEY, snapshot)
self._published_version = version
def tick(self) -> bool:
"""Publish if due. True if a snapshot was written."""
"""Apply a pending clear and publish if due. True if a snapshot was written."""
with self._tick_lock:
try:
cleared = self._apply_clear_request()
version = self.aggregator.version
now = self._clock()
if version == self._published_version:
return False
if (self._last_attempt is not None
and now - self._last_attempt < self.min_interval):
return False
if not cleared:
if version == self._published_version:
return False
if (self._last_attempt is not None
and now - self._last_attempt < self.min_interval):
return False
self._publish(version, now)
return True
except Exception as err: # never let reporting break the display
@@ -711,13 +754,16 @@ def apply_error_clear(request_id: str, args: Any) -> Dict[str, Any]:
# --- Reading side (web interface) -------------------------------------------
def read_error_report(cache_manager: Any) -> Optional[Dict[str, Any]]:
"""The display service's latest snapshot, or None before it has published.
def read_error_report(cache_manager: Any) -> Tuple[Optional[Dict[str, Any]], Optional[Dict[str, Any]]]:
"""The display service's latest snapshot and the latest clear request.
memory_ttl=0: the display writes the key, so only the file is current.
memory_ttl=0: both keys are written by the other process, so only the
file is current.
"""
snapshot = cache_manager.get(ERROR_SNAPSHOT_KEY, max_age=None, memory_ttl=0)
return snapshot if isinstance(snapshot, dict) else None
clear_request = cache_manager.get(ERROR_CLEAR_REQUEST_KEY, max_age=None, memory_ttl=0)
return (snapshot if isinstance(snapshot, dict) else None,
clear_request if isinstance(clear_request, dict) else None)
def _epoch(iso: Any) -> Optional[float]:
@@ -730,6 +776,36 @@ def _epoch(iso: Any) -> Optional[float]:
return None
def _pending_cutoff(snapshot: Optional[Dict[str, Any]],
clear_request: Optional[Dict[str, Any]]) -> Optional[float]:
"""The cutoff of a clear the snapshot has not applied yet, if any."""
if not clear_request:
return None
request_id = clear_request.get("request_id")
if not request_id:
return None
if snapshot is not None and snapshot.get("applied_clear_id") == request_id:
return None
try:
cutoff = float(clear_request.get("cutoff"))
except (TypeError, ValueError):
return None
if not math.isfinite(cutoff):
return None
# A wider clear has been applied since (over the control socket): this
# older request has nothing left to hide.
applied = snapshot.get("applied_clear_cutoff") if snapshot is not None else None
if (isinstance(applied, (int, float)) and not isinstance(applied, bool)
and applied >= cutoff):
return None
return cutoff
def _is_after(item: Any, field_name: str, cutoff: float) -> bool:
when = _epoch(item.get(field_name)) if isinstance(item, dict) else None
return when is not None and when > cutoff
def _empty_summary(snapshot: Optional[Dict[str, Any]]) -> Dict[str, Any]:
return {
"session_start": snapshot.get("session_start") if snapshot else None,
@@ -742,11 +818,11 @@ def _empty_summary(snapshot: Optional[Dict[str, Any]]) -> Dict[str, Any]:
}
def error_summary_from_report(snapshot: Optional[Dict[str, Any]]) -> Dict[str, Any]:
def error_summary_from_report(snapshot: Optional[Dict[str, Any]],
clear_request: Optional[Dict[str, Any]]) -> Dict[str, Any]:
"""The /errors/summary payload: get_error_summary()'s shape plus
``generated_at``, ``snapshot_available`` and ``clear_pending`` (always
False now: a clear is applied before its route answers; kept for API
compatibility)."""
``generated_at``, ``snapshot_available`` and ``clear_pending``."""
cutoff = _pending_cutoff(snapshot, clear_request)
summary = _empty_summary(snapshot)
if snapshot is not None:
for name, default in summary.items():
@@ -755,17 +831,32 @@ def error_summary_from_report(snapshot: Optional[Dict[str, Any]]) -> Dict[str, A
isinstance(default, float) and isinstance(value, int)) or (
name == "session_start" and isinstance(value, str)):
summary[name] = value
if cutoff is not None:
recent = summary["recent_errors"]
newest = _epoch(recent[-1].get("timestamp")) if recent and isinstance(recent[-1], dict) else None
if newest is None or newest <= cutoff:
# Everything the display has reported predates the clear.
summary = _empty_summary(snapshot)
else:
# Only part of it does. The lists can be filtered exactly; the
# counts cannot, and stay as reported until the display
# applies the clear (clear_pending says so).
summary["recent_errors"] = [r for r in recent if _is_after(r, "timestamp", cutoff)]
summary["active_patterns"] = {
k: p for k, p in summary["active_patterns"].items()
if _is_after(p, "last_seen", cutoff)
}
summary["generated_at"] = snapshot.get("generated_at") if snapshot else None
summary["snapshot_available"] = snapshot is not None
summary["clear_pending"] = False
summary["clear_pending"] = cutoff is not None
return summary
def plugin_health_from_report(snapshot: Optional[Dict[str, Any]],
clear_request: Optional[Dict[str, Any]],
plugin_id: str) -> Dict[str, Any]:
"""The /errors/plugin/<id> payload: get_plugin_health()'s shape plus
``generated_at``, ``snapshot_available`` and ``clear_pending`` (always
False, as in error_summary_from_report)."""
``generated_at``, ``snapshot_available`` and ``clear_pending``."""
health: Dict[str, Any] = {
"plugin_id": plugin_id,
"status": "healthy",
@@ -774,15 +865,19 @@ def plugin_health_from_report(snapshot: Optional[Dict[str, Any]],
"recent_error_count": 0,
"last_error": None,
}
cutoff = _pending_cutoff(snapshot, clear_request)
table = snapshot.get("plugin_health") if snapshot else None
entry = table.get(plugin_id) if isinstance(table, dict) else None
if isinstance(entry, dict):
for name in ("status", "total_errors", "error_types", "recent_error_count", "last_error"):
if name in entry:
health[name] = entry[name]
# last_error is the plugin's newest error: if even that predates a
# pending clear, so does everything else the display reported for it.
if cutoff is None or _is_after(entry.get("last_error"), "timestamp", cutoff):
for name in ("status", "total_errors", "error_types", "recent_error_count", "last_error"):
if name in entry:
health[name] = entry[name]
health["generated_at"] = snapshot.get("generated_at") if snapshot else None
health["snapshot_available"] = snapshot is not None
health["clear_pending"] = False
health["clear_pending"] = cutoff is not None
return health
@@ -804,35 +899,61 @@ def _count_cleared(summary: Dict[str, Any], cutoff: float) -> Optional[int]:
#: ``send(request_id, cutoff)`` hands a clear to the display over the control
#: socket and returns its ErrorsClearResult. It raises when the socket could
#: not carry it or the display failed it (``src.ipc.client.ControlError``).
ClearSender = Callable[[str, float], Dict[str, Any]]
#: socket and returns its ErrorsClearResult, or None when the socket could
#: not carry it and the mailbox should be written instead. Any exception it
#: raises reaches the caller: the display had the request and failed it.
ClearSender = Callable[[str, float], Optional[Dict[str, Any]]]
def request_error_clear(cache_manager: Any, cutoff: float,
send: ClearSender) -> Dict[str, Any]:
send: Optional[ClearSender] = None) -> Dict[str, Any]:
"""Ask the display service to forget errors recorded at or before ``cutoff``.
Over the control socket: the display applies the clear and republishes
its snapshot before it answers, so nothing is written here. Whatever
``send`` raises reaches the caller.
Over the control socket when ``send`` is given and carries it: the
display applies the clear and republishes its snapshot before it
answers, so nothing is written here. Otherwise (no socket, or a display
older than ``errors.clear``) a request is written to the
``plugin_error_clear_request`` mailbox, which the display applies on
its next tick, and readers hide the cleared errors until then.
Returns ``request_id``, ``cutoff`` (ISO, local time), ``cleared_count``
(the display's own count, else an estimate from the snapshot; see
_count_cleared), and ``clear_requested``, ``applied`` and ``transport``,
which are always True, True and ``"socket"`` now and are kept for API
compatibility.
Returns ``request_id``, ``cutoff`` (ISO, local time), ``cleared_count``,
``clear_requested``, ``applied`` (the display has already cleared them)
and ``transport`` (``socket`` or ``mailbox``). ``cleared_count`` is the
display's own count over the socket, else an estimate from the snapshot
(see _count_cleared). Raises OSError when a mailbox request did not reach
the shared cache, since a cache without a usable directory accepts set()
and keeps nothing.
A request the display has not applied yet is only ever widened: a later,
narrower one ("older than 24 hours" after "everything") overwriting it
would otherwise bring back the errors the first one hid.
"""
before = error_summary_from_report(read_error_report(cache_manager))
snapshot, clear_request = read_error_report(cache_manager)
pending = _pending_cutoff(snapshot, clear_request)
if pending is not None:
cutoff = max(cutoff, pending)
before = error_summary_from_report(snapshot, clear_request)
request_id = uuid.uuid4().hex
result = send(request_id, cutoff)
count = result.get("cleared") if isinstance(result, dict) else None
return {
answer = {
"clear_requested": True,
"request_id": request_id,
"cutoff": datetime.fromtimestamp(cutoff).isoformat(),
"applied": True,
"transport": "socket",
"cleared_count": (count if isinstance(count, int) and not isinstance(count, bool)
else _count_cleared(before, cutoff)),
}
if send is not None:
result = send(request_id, cutoff)
if result is not None:
count = result.get("cleared")
return dict(answer, applied=True, transport="socket",
cleared_count=count if isinstance(count, int) and not isinstance(count, bool)
else _count_cleared(before, cutoff))
request = {
"request_id": request_id,
"cutoff": cutoff,
"requested_at": time.time(),
}
cache_manager.set(ERROR_CLEAR_REQUEST_KEY, request)
stored = cache_manager.get(ERROR_CLEAR_REQUEST_KEY, max_age=None, memory_ttl=0)
if not isinstance(stored, dict) or stored.get("request_id") != request["request_id"]:
raise OSError("the clear request was not stored in the shared cache")
return dict(answer, applied=False, transport="mailbox",
cleared_count=_count_cleared(before, cutoff))
+28 -23
View File
@@ -5,11 +5,11 @@ a refused or timed-out connection, a reply that breaks the contract, or an
error the display returned -- raises :class:`ControlError` with a short
``reason``. Nothing here blocks for longer than ``timeout`` in total.
There is no other way to reach the display: the file mailboxes the web
interface used to fall back to are gone. :func:`display_not_listening` tells
a caller when no display is listening yet (it is stopped, or still
starting), the one case where sending the same request again later, once a
display is up, can work.
Whether the caller may then write the file mailbox instead is
:func:`should_fall_back`: only when the display never took the request (it
could not be reached, or it is too old to know the command). A display that
took the request and then failed, refused or went quiet is answered as
that, not posted a second time through the mailbox.
"""
from __future__ import annotations
@@ -43,7 +43,8 @@ from src.ipc.contract import (
#: Total budget for one request: connect, send and the reply. The display
#: answers from a thread that does no rendering, normally within a few
#: milliseconds; this only bounds a wedged one.
#: milliseconds; this only bounds a wedged one. The web route then falls back
#: to the mailbox, so a timeout costs this much latency and nothing else.
DEFAULT_TIMEOUT_SECONDS = 1.0
@@ -73,25 +74,28 @@ class ControlError(Exception):
#: Answers from a display that read the request but does not speak it: one
#: older than the command (an upgrade in progress) or the protocol version.
#: It did nothing, so the mailbox is the way to reach it.
UPGRADE_REASONS = frozenset({ErrorCode.UNKNOWN_COMMAND, ErrorCode.UNSUPPORTED_VERSION})
#: Transport reasons that mean nothing is listening at the socket: no socket
#: file (the display is stopped, or has not reached its run loop), or a file
#: nobody accepts on (a stale socket) or that this user may not open.
NOT_LISTENING_REASONS = frozenset({'no_socket', 'refused'})
def should_fall_back(error: BaseException) -> bool:
"""May the caller write the file mailbox after ``error``?
def display_not_listening(error: BaseException) -> bool:
"""True when ``error`` says no display took the request because none is
listening: it never reached one (``sent`` is False) and the reason is in
:data:`NOT_LISTENING_REASONS`. A display that is started, or finishes
starting, may take the same request later. False for everything else:
a display that had the request and failed it, one too old to know the
command, a client that cannot use the socket at all (``disabled``,
``unsupported``), and an exception that is not a :class:`ControlError`.
Yes when the display never took the request: there is no socket (the
display is stopped, predates the socket, or it is switched off), the
connection was refused or timed out, the display turned the connection
away before reading it, or it is too old to know the command
(:data:`UPGRADE_REASONS`). Also for an error that is not a
:class:`ControlError` (a bug in the client), as before.
No once the display had the request: a ``busy`` queue, ``invalid_args``,
an ``internal`` error, or a timeout or hang-up after the request was
sent. The display may have applied it, or would refuse it from the
mailbox too, so a second copy there only hides the failure.
"""
return (isinstance(error, ControlError) and not error.sent
and error.reason in NOT_LISTENING_REASONS)
if not isinstance(error, ControlError):
return True
return not error.sent or error.reason in UPGRADE_REASONS
def request(cmd: str, args: Optional[Mapping[str, Any]] = None, *,
@@ -214,8 +218,8 @@ def on_demand_start(request_id: str, plugin_id: Optional[str], mode: Optional[st
paths: Optional[Sequence[str]] = None) -> Dict[str, Any]:
"""Ask the display to show a plugin now. Returns the ack; raises :class:`ControlError`.
``request_id`` doubles as the on-demand request id, so a request sent
twice with the same id is processed only once.
``request_id`` doubles as the on-demand request id, so a request that a
timed-out caller then also writes to the mailbox is processed only once.
"""
args = {'plugin_id': plugin_id, 'mode': mode, 'duration': duration, 'pinned': pinned}
return request(Command.ON_DEMAND_START, args, request_id=request_id,
@@ -277,7 +281,8 @@ def errors_clear(request_id: str, cutoff: float, *,
Returns :class:`~src.ipc.contract.ErrorsClearResult` once it is done.
Raises :class:`ControlError`: ``unknown_command`` from a display older
than the command.
than the command, which still reads the ``plugin_error_clear_request``
mailbox.
"""
return request(Command.ERRORS_CLEAR, {'cutoff': cutoff}, request_id=request_id,
timeout=timeout, paths=paths)
+11 -10
View File
@@ -85,8 +85,8 @@ DEFAULT_SOCKET_PATH = DEFAULT_SOCKET_DIR + '/' + SOCKET_NAME
#: Overrides the socket path for both processes (a dev checkout, a second
#: instance, tests). One of :data:`DISABLED_VALUES` turns the socket off: the
#: display does not serve it, and the web interface cannot send it commands
#: (it still reads the state the display writes to the cache).
#: display does not serve it and the web interface goes straight to the
#: file mailbox.
SOCKET_PATH_ENV = 'LEDMATRIX_CONTROL_SOCKET'
DISABLED_VALUES = frozenset({'off', '0', 'false', 'no', 'none', 'disabled'})
@@ -375,8 +375,8 @@ def _optional_name(args: Mapping[str, Any], key: str) -> Optional[str]:
def _optional_duration(value: Any) -> Optional[float]:
"""Seconds, or None for "until stopped". 0 means the same as None.
Numbers and numeric strings are accepted, the same as the REST route
takes them; anything else is refused rather than guessed.
Numbers and numeric strings are accepted, the same as the REST route and
the file mailbox take them; anything else is refused rather than guessed.
"""
if value is None or value == '':
return None
@@ -418,7 +418,7 @@ class HelloArgs:
class OnDemandStartArgs:
"""``on_demand.start``: show a plugin (or one of its modes) now.
The same fields as the REST route's body. At least one of ``plugin_id``
The same fields the file mailbox carries. At least one of ``plugin_id``
and ``mode`` is required; the display resolves the other.
"""
plugin_id: Optional[str] = None
@@ -628,12 +628,13 @@ def parse_args(cmd: str, args: Mapping[str, Any]) -> CommandArgs:
def on_demand_request(request_id: str, args: Union[OnDemandStartArgs, OnDemandStopArgs],
timestamp: float) -> Dict[str, Any]:
"""The on-demand request dict for a queued on-demand command.
"""The file-mailbox payload for a queued on-demand command.
The display hands socket commands to the same code that handles
plugins' own requests (``DisplayController._handle_on_demand_request``),
so a command behaves identically whichever way it arrived. (This was
the file mailbox's payload, which the display no longer reads.)
The display hands socket commands to the same code that handles the
mailbox (``DisplayController._handle_on_demand_request``), so a command
behaves identically whichever way it arrived, and a request that came
both ways (a client that timed out and fell back) is processed once: the
request id is the same.
"""
if isinstance(args, OnDemandStartArgs):
return {'request_id': request_id, 'action': 'start', 'plugin_id': args.plugin_id,
+12 -13
View File
@@ -3,9 +3,9 @@
A small threaded server on a Unix stream socket (``/run/ledmatrix/control.sock``
by default; see :mod:`src.ipc.contract` for the protocol). It never touches
rendering: a command that changes the panel is validated, put on a bounded
queue and acknowledged, and the render thread drains that queue
(``DisplayController._poll_on_demand_requests``), handing each on-demand
command to the code that handles plugins' own requests. Queries (``on_demand.status``) are
queue and acknowledged, and the render thread drains that queue at the point
where it reads the file mailbox (``DisplayController._poll_on_demand_requests``),
handing each command to the same code. Queries (``on_demand.status``) are
answered from a snapshot callable the display provides, and the few commands
that touch nothing the render thread owns (``errors.clear``) by a handler the
display registers, on the connection thread.
@@ -110,8 +110,8 @@ MAX_CLIENTS = 8
#: Commands waiting for the render thread. It drains them at least every
#: 0.25 s, so a full queue means the render thread is stuck, and the client
#: is told ``busy`` instead of piling up work, and the web interface reports
#: the failure.
#: is told ``busy`` instead of piling up work. The mailbox would not be read
#: either, so the web interface reports the failure rather than fall back.
QUEUE_SIZE = 16
#: Timeout for one recv()/send() on a connection.
@@ -185,7 +185,7 @@ class QueuedCommand:
outcome: Optional[CommandOutcome] = field(default=None, compare=False, repr=False)
def as_on_demand_request(self) -> Dict[str, Any]:
"""The on-demand request dict the display's on-demand handler takes."""
"""The mailbox-shaped payload the display's on-demand handler takes."""
if not isinstance(self.args, (OnDemandStartArgs, OnDemandStopArgs)):
raise TypeError(f'{self.cmd} is not an on-demand command')
return on_demand_request(self.request_id, self.args, self.received_at)
@@ -619,8 +619,8 @@ class ControlServer:
def start(self) -> bool:
"""Bind and start serving. False (logged) when the socket cannot be served.
Never raises: without the socket the display runs, but the web
interface cannot send it commands.
Never raises: without the socket the web interface uses the file
mailbox, exactly as before.
"""
if not socket_supported():
logger.debug("Control socket not started: no Unix sockets on this platform")
@@ -632,7 +632,7 @@ class ControlServer:
self._bind()
except OSError as e:
logger.warning("Control socket not started at %s (%s); the web interface "
"cannot send this display commands", self.path, e)
"will use the file mailbox", self.path, e)
self._close_socket()
return False
self._stopping.clear()
@@ -1134,13 +1134,12 @@ def start_control_server(status_provider: Optional[StatusProvider] = None,
"""Start the display's control socket, or return None when it can't run.
None covers Windows, ``LEDMATRIX_CONTROL_SOCKET=off`` and any failure to
bind; in every case the web interface cannot send the display commands,
and reads the cache keys the display still writes.
bind; in every case the web interface falls back to the file mailbox
and to the cache keys the display still writes.
"""
path = server_socket_path(environ)
if path is None:
logger.debug("Control socket disabled or unsupported here; the web interface "
"cannot send this display commands")
logger.debug("Control socket disabled or unsupported here; using the file mailbox only")
return None
server = ControlServer(path, status_provider, resolve_socket_group(cache_dir),
state_hub=state_hub, handlers=handlers)
+7 -7
View File
@@ -1093,16 +1093,16 @@ class BasePlugin(ABC):
Returns:
The request id once the display has queued it, or None when
there is no display in this process to ask (the web interface,
scripts/check_plugin.py) or its queue is full. Nothing else can
take the request then: the ``display_on_demand_request`` file
mailbox older cores read is gone, and a write to it is dropped
with a warning. See "On-demand display" in
docs/PLUGIN_API_REFERENCE.md.
scripts/check_plugin.py) or its queue is full. A plugin that
also runs on cores without this method writes the
``display_on_demand_request`` mailbox on None, as before; see
"On-demand display" in docs/PLUGIN_API_REFERENCE.md.
Example::
if self.request_on_demand(mode='my_alert', duration=15) is None:
self.logger.info("No display to show the alert on")
if not (hasattr(self, 'request_on_demand')
and self.request_on_demand(mode='my_alert', duration=15)):
self._write_on_demand_mailbox(...) # older cores
"""
request = getattr(getattr(self, 'plugin_manager', None), 'request_on_demand', None)
if not callable(request):
+1 -1
View File
@@ -1863,7 +1863,7 @@ class PluginManager:
"""Route plugins' on-demand requests to ``handler`` (None: nowhere).
The display controller sets its ``submit_plugin_on_demand`` here
before any plugin loads. The handler takes an on-demand request dict
before any plugin loads. The handler takes a mailbox-shaped request
from any thread, queues it for the render thread and returns True,
or False when it could not. A plugin manager with no handler (the
web interface's, a test's, scripts/check_plugin.py's) has no screen
+15 -11
View File
@@ -4,6 +4,7 @@ Centralized error handling for web interface.
Provides helpers for consistent error responses across API endpoints.
"""
import errno
from typing import Any, Optional
from flask import jsonify
@@ -20,10 +21,9 @@ logger = get_logger(__name__)
_MAX_DETAIL_LENGTH = 400
def describe_exception(exc: BaseException,
max_length: int = _MAX_DETAIL_LENGTH) -> str:
def describe_exception(exc: BaseException) -> str:
"""
One-line, safe-to-return description of an exception.
Machine-readable reason code for an exception, safe to return over HTTP.
The generic "an error occurred; see logs for details" tells a user nothing
and, when the failure is bad enough, the logs are unreachable too: a device
@@ -31,20 +31,24 @@ def describe_exception(exc: BaseException,
*including* the log viewer, because journalctl could not be executed. The
underlying `[Errno 5] Input/output error` named the fault immediately.
Returns "TypeName: message", credentials redacted and length capped. The
type alone is worth carrying -- a bare PermissionError says more than any
generic sentence.
So the type and errno still go back -- "OSError:EIO", "PermissionError:
EACCES", "TimeoutExpired" -- but never the exception's message, which can
quote paths, URLs, credentials or a library's internals (CodeQL
py/stack-trace-exposure). The message is logged here instead, so every
reason code a client sees has its full text in the log.
Args:
exc: The exception to describe
max_length: Truncate beyond this many characters
Returns:
A single-line description, never empty
"TypeName" or "TypeName:ERRNO", never empty
"""
message = str(exc).strip()
text = f"{type(exc).__name__}: {message}" if message else type(exc).__name__
return redact_text(text, max_length)
code = type(exc).__name__
exc_errno = getattr(exc, 'errno', None)
if isinstance(exc_errno, int) and exc_errno in errno.errorcode:
code = f"{code}:{errno.errorcode[exc_errno]}"
logger.warning("Error reported to the client as %s: %s", code, redact_text(str(exc)))
return code
def redact_text(text: str, max_length: int = _MAX_DETAIL_LENGTH) -> str:
+9 -9
View File
@@ -1369,7 +1369,7 @@ class WiFiManager:
self.enable_ap_mode(force=True)
except Exception as ap_error: # nosec B110 - last-resort; do not re-raise, but log for debugging
logger.error("Last-resort AP mode enable failed in recovery path: %s", ap_error, exc_info=True)
return False, str(e)
return False, f"Connection failed ({type(e).__name__}); see logs for details"
def _failsafe_ap(self, enabled_msg: str, failed_msg: str) -> Tuple[bool, str]:
"""Force the setup AP up after a connect that left no working network,
@@ -1585,7 +1585,7 @@ class WiFiManager:
except Exception as e:
logger.error(f"Error connecting with nmcli: {e}")
self._show_led_message("Connection error", duration=5)
return False, str(e)
return False, f"Connection failed ({type(e).__name__}); see logs for details"
# 802.11 caps an SSID at 32 octets. Control characters cannot appear in a
# real one, and a leading "-" would be read by nmcli as an option rather
@@ -1725,7 +1725,7 @@ class WiFiManager:
return False, "nmcli is required to disconnect from WiFi"
except Exception as e:
logger.error(f"Error disconnecting from WiFi: {e}")
return False, str(e)
return False, f"Disconnect failed ({type(e).__name__}); see logs for details"
def _ensure_wifi_radio_enabled(self, max_retries: int = 3) -> bool:
"""
@@ -2004,7 +2004,7 @@ class WiFiManager:
return False, "No WiFi tools available (nmcli, hostapd, or dnsmasq required)"
except Exception as e:
logger.error(f"Error in enable_ap_mode: {e}")
return False, str(e)
return False, f"Could not enable AP mode ({type(e).__name__}); see logs for details"
def _mark_forced(self) -> None:
"""Record that AP mode was forced on, so the periodic check leaves it
@@ -2099,10 +2099,10 @@ class WiFiManager:
return True, "AP mode enabled"
except Exception as e:
logger.error(f"Error starting AP services: {e}")
return False, str(e)
return False, f"Could not enable AP mode ({type(e).__name__}); see logs for details"
except Exception as e:
logger.error(f"Error enabling AP mode: {e}")
return False, str(e)
return False, f"Could not enable AP mode ({type(e).__name__}); see logs for details"
def _enable_ap_mode_nmcli_hotspot(self) -> Tuple[bool, str]:
"""
@@ -2227,7 +2227,7 @@ class WiFiManager:
logger.error(f"Error starting AP mode with nmcli: {e}")
self._remove_nm_dnsmasq_captive_conf()
self._show_led_message("Setup mode error", duration=5)
return False, str(e)
return False, f"Could not enable AP mode ({type(e).__name__}); see logs for details"
def _get_ap_status_nmcli(self) -> Dict:
"""
@@ -2409,10 +2409,10 @@ class WiFiManager:
return True, "AP mode disabled"
except Exception as e:
logger.error(f"Error stopping AP services: {e}")
return False, str(e)
return False, f"Could not disable AP mode ({type(e).__name__}); see logs for details"
except Exception as e:
logger.error(f"Error disabling AP mode: {e}")
return False, str(e)
return False, f"Could not disable AP mode ({type(e).__name__}); see logs for details"
def _create_hostapd_config(self):
"""Create hostapd configuration file"""
+9 -15
View File
@@ -155,6 +155,13 @@ class FakeCache:
self._writes += 1
self._written[key] = self._writes
def file_signature(self, key):
"""CacheManager.file_signature: None without a file, else a value
that changes with every write."""
if key not in self.data:
return None
return (self._written.get(key, 0), 0, 0)
def delete(self, key):
self.data.pop(key, None)
@@ -722,23 +729,10 @@ class RunLoopHarness:
self.controller.available_modes.append(mode)
def on_demand_request(self, t: float, request_id: str, action: str = "start", **fields):
"""An on-demand start or stop from the web interface at ``t``: a
command on the control socket (served for the run if no test did),
which is the only way the web interface reaches the display."""
from src.ipc.contract import Command, parse_args
from src.ipc.server import QueuedCommand
server = self.controller._control_server
if not isinstance(server, FakeControlServer):
server = self.control_socket()
cmd = Command.ON_DEMAND_START if action == "start" else Command.ON_DEMAND_STOP
command = QueuedCommand(request_id=request_id, cmd=cmd,
args=parse_args(cmd, fields if action == "start" else {}),
received_at=0.0)
def post():
self.log("request", f"{action}:{request_id}")
server.queue.append(command)
self.cache.set("display_on_demand_request",
{"request_id": request_id, "action": action, **fields})
self.clock.at(t, post)
def restore_on_demand(self, plugin_id: str, mode: Optional[str] = None,
-14
View File
@@ -328,20 +328,6 @@ def _hermetic_control_socket(monkeypatch):
monkeypatch.setenv(SOCKET_PATH_ENV, 'off')
@pytest.fixture(autouse=True)
def _no_pending_on_demand_dispatch():
"""Drop the web process's on-demand dispatcher after each test, so a
start one test left pending is not still being sent in the next."""
yield
module = sys.modules.get('web_interface.on_demand_dispatch')
if module is None:
return
dispatcher = module.current()
if dispatcher is not None:
dispatcher.cancel('test-teardown')
module.reset_for_tests()
@pytest.fixture(autouse=True)
def _hermetic_unit_refresh(monkeypatch, tmp_path_factory):
"""Keep updates' systemd unit refresh off the host.
+3 -3
View File
@@ -1,16 +1,16 @@
{
"screens": [
[0.0, "clock", 20.0, "duration", 20, false],
[20.0, "weather", 5.0, "on-demand-start", 5, true],
[20.0, "weather", 5.0, "on-demand-start", 6, true],
[25.0, "sports_recent", 15.0, "duration", 15, true],
[40.0, "sports_upcoming", 15.0, "duration", 15, true],
[55.0, "sports_recent", 15.0, "duration", 15, true],
[70.0, "sports_upcoming", 15.0, "duration", 15, true],
[85.0, "sports_recent", 10.0, "on-demand-requested-stop", 10, true],
[85.0, "sports_recent", 10.0, "on-demand-requested-stop", 11, true],
[95.0, "weather", 20.0, "duration", 20, true],
[115.0, "sports_recent", 15.0, "duration", 15, true],
[130.0, "sports_upcoming", 15.0, "duration", 15, true],
[145.0, "clock", 5.0, "on-demand-start", 5, true],
[145.0, "clock", 5.0, "on-demand-start", 6, true],
[150.0, "weather", 20.0, "duration", 20, true],
[170.0, "weather", 10.0, "on-demand-expired", 10, true],
[180.0, "clock", 20.0, "duration", 20, true],
+4 -4
View File
@@ -1,18 +1,18 @@
{
"screens": [
[0.0, "clock", 5.0, "on-demand-start", 5, false],
[0.0, "clock", 5.0, "on-demand-start", 6, false],
[5.0, "sports_live", 15.0, "duration", 15, true],
[20.0, "sports_recent", 15.0, "duration", 15, true],
[35.0, "sports_upcoming", 5.0, "on-demand-requested-stop", 5, true],
[35.0, "sports_upcoming", 5.0, "on-demand-requested-stop", 6, true],
[40.0, "clock", 20.0, "duration", 20, true],
[60.0, "sports_live", 15.0, "display-false", 11, true],
[75.0, "sports_recent", 15.0, "duration", 15, true],
[90.0, "sports_upcoming", 10.0, "on-demand-start", 10, true],
[90.0, "sports_upcoming", 10.0, "on-demand-start", 11, true],
[100.0, "sports_live", 0.0, "empty", 1, true],
[100.0, "sports_recent", 15.0, "duration", 15, true],
[115.0, "sports_upcoming", 15.0, "duration", 15, true],
[130.0, "sports_live", 0.0, "empty", 1, true],
[130.0, "sports_recent", 10.0, "on-demand-requested-stop", 10, true],
[130.0, "sports_recent", 10.0, "on-demand-requested-stop", 11, true],
[140.0, "sports_upcoming", 15.0, "duration", 15, true],
[155.0, "clock", 5.0, "horizon", 5, true]
],
+2 -2
View File
@@ -1,11 +1,11 @@
{
"screens": [
[0.0, "clock", 12.0, "on-demand-start", 12, false],
[0.0, "clock", 12.0, "on-demand-start", 13, false],
[12.0, "sports_upcoming", 15.0, "duration", 15, true],
[27.0, "sports_upcoming", 15.0, "duration", 15, true],
[42.0, "sports_upcoming", 15.0, "duration", 15, true],
[57.0, "sports_upcoming", 15.0, "duration", 15, true],
[72.0, "sports_upcoming", 8.0, "on-demand-start", 8, true],
[72.0, "sports_upcoming", 8.0, "on-demand-start", 9, true],
[80.0, "app_a", 0.0, "empty", 1, true],
[80.0, "app_b", 10.0, "duration", 10, true],
[90.0, "app_a", 0.0, "empty", 1, true],
+11 -11
View File
@@ -6,22 +6,22 @@
[70.255, "sports_live", 20.0, "duration", 20, true],
[90.255, "sports_live", 20.0, "display-false", 11, false],
[110.255, "<vegas>", 30.008, "duration", 3751, null],
[140.263, "<vegas>", 9.744, "on-demand-start", 1218, null],
[150.007, "clock", 20.0, "duration", 20, true],
[170.007, "clock", 5.0, "on-demand-expired", 5, true],
[175.007, "<vegas>", 25.999, "vegas-interrupt", 3250, null],
[201.006, "<wifi>", 2.0, "duration", 4, null],
[203.006, "<vegas>", 30.008, "duration", 3751, null],
[233.014, "<vegas>", 26.986, "horizon", 3374, null]
[140.263, "<vegas>", 10.0, "on-demand-start", 1250, null],
[150.263, "clock", 20.0, "duration", 20, true],
[170.263, "clock", 5.0, "on-demand-expired", 5, true],
[175.263, "<vegas>", 24.959, "vegas-interrupt", 3120, null],
[200.222, "<wifi>", 3.0, "duration", 6, null],
[203.222, "<vegas>", 30.008, "duration", 3751, null],
[233.23, "<vegas>", 26.77, "horizon", 3347, null]
],
"events": [
[70.255, "vegas-live"],
[70.255, "live", "sports_live"],
[150.0, "request", "start:v1"],
[150.007, "on-demand-start", "clock"],
[150.007, "vegas-interrupt"],
[175.007, "on-demand-expired"],
[150.263, "on-demand-start", "clock"],
[150.263, "vegas-interrupt"],
[175.263, "on-demand-expired"],
[200.0, "wifi-file", "Connected to HomeNet"],
[201.006, "vegas-interrupt"]
[200.222, "vegas-interrupt"]
]
}
+1 -2
View File
@@ -32,8 +32,7 @@ const UNIT = ['unit/test_list_filter.js', 'unit/test_render_cards.js',
'unit/test_page_registry.js', 'unit/test_core_modules.js',
'unit/test_overview_reconciliation_poll.js',
'unit/test_display_partial_ids.js',
'unit/test_general_web_login_token.js',
'unit/test_on_demand_starting.js'];
'unit/test_general_web_login_token.js'];
const DOM = ['dom/test_installed_dom.js', 'dom/test_store_dom.js', 'dom/test_no_double_fetch.js',
'dom/test_tools_sections.js', 'dom/test_cache_page.js',
'dom/test_durations_page.js', 'dom/test_operation_history_page.js',
-69
View File
@@ -1,69 +0,0 @@
// POST /api/v3/display/on-demand/start answers 202 with status "starting"
// when the display service has to be started first: the request is taken,
// and the web process sends it once the display listens. "Preview on
// display" (app.js) must read that as taken -- an info toast and the
// floating preview opened -- not as a failure. Runs the shipped app.js in a
// vm with a minimal fake DOM, as test_restart_banner.js does.
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const V3 = path.resolve(__dirname, '../../../web_interface/static/v3');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' ' + JSON.stringify(extra) : '')));
function load(answer) {
const notes = [];
const opened = [];
const noop = () => {};
const document = {
body: { addEventListener: noop },
addEventListener: noop,
getElementById: () => null,
querySelector: () => null,
querySelectorAll: () => [],
};
const window = { addEventListener: noop, getApp: () => null };
const context = {
window, document, console,
sessionStorage: { setItem: noop, removeItem: noop, getItem: () => null },
showNotification: (m, t) => notes.push([m, t]),
setTimeout: () => 0,
fetch: () => Promise.resolve({ json: () => Promise.resolve(answer) }),
};
vm.createContext(context);
vm.runInContext(fs.readFileSync(path.join(V3, 'app.js'), 'utf8'), context);
window.toggleFloatingPreview = (open) => opened.push(open);
return { window, notes, opened };
}
async function preview(answer) {
const t = load(answer);
t.window.previewPluginNow('weather');
for (let i = 0; i < 5; i++) await Promise.resolve();
return t;
}
(async () => {
console.log('\npreviewPluginNow');
{
const t = await preview({ status: 'starting', message: 'The display service is starting',
data: { request_id: 'r1', pending: true } });
ok('a 202 "starting" answer is an info toast, not an error',
t.notes.length === 1 && t.notes[0][1] === 'info', t.notes);
ok('and the preview opens', t.opened.length === 1 && t.opened[0] === true, t.opened);
}
{
const t = await preview({ status: 'success', data: { request_id: 'r1' } });
ok('a 200 success still opens it', t.opened.length === 1 && t.notes[0][1] === 'success', t.notes);
}
{
const t = await preview({ status: 'error', message: 'no display' });
ok('an error does not', t.opened.length === 0 && t.notes[0][1] === 'error', t.notes);
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})();
+1 -4
View File
@@ -54,10 +54,7 @@ class TestOnDemandStart:
def service(self, api_v3_module):
api_v3_module.api_v3.plugin_catalog = None
api_v3_module.api_v3.config_manager = None
# The display answers the control socket (the only way in).
with patch("web_interface.blueprints.api_v3.display.control_client.on_demand_start",
side_effect=lambda request_id, *a: {"accepted": True}), \
patch("web_interface.blueprints.api_v3.display._get_display_service_status",
with patch("web_interface.blueprints.api_v3.display._get_display_service_status",
return_value={"active": True}), \
patch("web_interface.blueprints.api_v3.display._stop_display_service") as stop, \
patch("web_interface.blueprints.api_v3.display._ensure_display_service_running",
+2 -1
View File
@@ -147,7 +147,8 @@ class TestOneBadConfigSectionDoesNotBlankTheList:
side_effect=RuntimeError("disk is gone"))
resp = api_v3_client.get('/api/v3/display/modes')
assert resp.status_code == 500
assert 'disk is gone' in resp.get_json()['details']
assert resp.get_json()['details'] == 'RuntimeError'
assert 'disk is gone' not in json.dumps(resp.get_json())
def test_credentials_in_the_exception_are_redacted(self, api_v3_module, api_v3_client):
"""describe_exception is what makes returning detail safe."""
+4 -15
View File
@@ -77,22 +77,11 @@ def fresh_web_process(api_v3_module, plugins_dir):
@pytest.fixture
def display_service(api_v3_module):
"""Keep on-demand start away from systemctl and the real cache; the
display's control socket is a mock that acks. What it was sent is
recorded as ``.set(cmd, request)`` calls on the yielded mock."""
api_v3_module.api_v3.cache_manager = MagicMock()
sent = MagicMock()
def ack(request_id, plugin_id, mode, *a, **kw):
sent.set('on_demand.start', {'request_id': request_id,
'plugin_id': plugin_id, 'mode': mode})
return {'accepted': True}
with patch('web_interface.blueprints.api_v3.display._get_display_service_status') as status, \
patch('web_interface.blueprints.api_v3.display.control_client.on_demand_start',
side_effect=ack):
"""Keep on-demand start away from systemctl and the real cache."""
cache = api_v3_module.api_v3.cache_manager = MagicMock()
with patch('web_interface.blueprints.api_v3.display._get_display_service_status') as status:
status.return_value = {'active': True}
yield sent
yield cache
def _start(client, **body):
+147
View File
@@ -0,0 +1,147 @@
"""No API response carries an exception's message (CodeQL py/stack-trace-exposure).
One representative route per file that had open alerts. Each forces a failure
whose message holds a marker and asserts the marker is nowhere in the body:
the message goes to the log, the client gets a fixed message plus a reason
code (describe_exception: the type, and the errno for an OSError).
"""
import json
import sys
from pathlib import Path
from unittest.mock import MagicMock, patch
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
LEAK = "LEAKED-/home/pi/secret token=abc123"
API = "web_interface.blueprints.api_v3"
def _assert_no_leak(response):
body = response.get_data(as_text=True)
assert "LEAKED" not in body, body
assert "abc123" not in body, body
return json.loads(body)
def test_display_service_status_drops_systemctl_output(api_v3_module, api_v3_client,
monkeypatch):
"""display.py: the on-demand routes return the service status verbatim."""
api_v3_module.api_v3.cache_manager.get.return_value = None
monkeypatch.setattr(f"{API}.display.display_state.read_state", lambda: None)
with patch(f"{API}.subprocess.run", side_effect=OSError(13, LEAK)):
body = _assert_no_leak(api_v3_client.get("/api/v3/display/on-demand/status"))
assert body["data"]["service"] == {"active": False, "returncode": -1}
@pytest.mark.parametrize("helper", ["_ensure_display_service_running",
"_stop_display_service"])
def test_service_results_keep_returncode_but_not_output(api_v3_module, helper):
"""display.py start/stop: returncode/active/started stay, stdout/stderr go."""
failed = MagicMock(returncode=1, stdout=LEAK, stderr=LEAK)
with patch(f"{API}.subprocess.run", return_value=failed):
result = getattr(api_v3_module, helper)()
assert "LEAKED" not in json.dumps(result)
assert result["returncode"] == 1 and result["active"] is False
assert "stdout" not in result and "stderr" not in result
def test_wifi_connect_failure(api_v3_client):
"""wifi.py: a raising connect, and the attempt /wifi/status reports after."""
with patch("src.wifi_manager.WiFiManager") as cls:
cls.return_value._is_ap_mode_active.return_value = False
cls.return_value.connect_to_network.side_effect = RuntimeError(LEAK)
body = _assert_no_leak(api_v3_client.post(
"/api/v3/wifi/connect", json={"ssid": "HomeNet", "password": "pw"}))
assert body["details"] == "RuntimeError"
cls.return_value.get_wifi_status.return_value = MagicMock(
connected=False, ssid=None, ip_address=None, signal=0, ap_mode_active=False)
cls.return_value.config = {}
status = _assert_no_leak(api_v3_client.get("/api/v3/wifi/status"))
assert status["data"]["last_connect_attempt"]["message"] == (
"Failed to connect to network (RuntimeError)")
def test_wifi_manager_messages_carry_no_exception_text():
"""src/wifi_manager.py: its (success, message) is what the wifi routes return."""
from src.wifi_manager import WiFiManager
manager = WiFiManager.__new__(WiFiManager) # no __init__: no host access
manager.get_wifi_status = MagicMock(side_effect=OSError(5, LEAK))
success, message = manager.disconnect_from_network()
assert success is False
assert "LEAKED" not in message and "OSError" in message
def test_system_action_exception(api_v3_client):
"""system.py: execute_system_action's catch-all."""
with patch("subprocess.run", side_effect=OSError(5, LEAK)):
body = _assert_no_leak(api_v3_client.post(
"/api/v3/system/action", json={"action": "stop_display"}))
assert body["details"] == "OSError:EIO"
def test_calendar_registration_failure(api_v3_client, tmp_path, monkeypatch):
"""plugin_calendar.py: the auth script could not be run."""
plugin_dir = tmp_path / "calendar"
plugin_dir.mkdir()
(plugin_dir / "credentials.json").write_text("{}", encoding="utf-8")
(plugin_dir / "calendar_registration.py").write_text("", encoding="utf-8")
monkeypatch.setattr(f"{API}._calendar_plugin_dir", lambda: plugin_dir)
with patch(f"{API}.subprocess.run", side_effect=OSError(13, LEAK)):
body = _assert_no_leak(api_v3_client.post(
"/api/v3/plugins/calendar/authenticate", json={"code": "x"}))
assert "EACCES" in body["message"]
def test_health_failure(api_v3_client, monkeypatch):
"""misc.py: get_health's catch-all."""
def boom():
raise RuntimeError(LEAK)
monkeypatch.setattr(f"{API}.misc._get_display_service_status", boom)
body = _assert_no_leak(api_v3_client.get("/api/v3/health"))
assert body["details"] == "RuntimeError"
def test_config_route_failure(api_v3_module, api_v3_client):
"""error_handler.py: create_error_response, as config.py's routes use it."""
api_v3_module.api_v3.config_manager.load_config.side_effect = RuntimeError(LEAK)
body = _assert_no_leak(api_v3_client.get("/api/v3/config/schedule"))
assert body["details"] == "RuntimeError"
def test_plugin_route_failure(api_v3_module, api_v3_client):
"""plugins.py: an unhandled error in a plugin route."""
api_v3_module.api_v3.plugin_catalog.get_all_plugin_info.side_effect = RuntimeError(LEAK)
body = _assert_no_leak(api_v3_client.get("/api/v3/plugins/installed"))
assert body["details"] == "RuntimeError"
def test_starlark_route_failure(api_v3_client):
"""starlark.py: one of its catch-alls."""
with patch(f"{API}._get_starlark_plugin", side_effect=RuntimeError(LEAK)):
body = _assert_no_leak(api_v3_client.get("/api/v3/starlark/status"))
assert body["details"] == "RuntimeError"
def test_unit_refresh_failure(monkeypatch):
"""system.py git_pull: perform_core_update appends unit_refresh's message."""
from web_interface import unit_refresh
def boom(*_a, **_k):
raise RuntimeError(LEAK)
monkeypatch.setattr(unit_refresh, "stale_units", boom)
result = unit_refresh.refresh_after_update()
assert result["status"] == unit_refresh.FAILED
assert "LEAKED" not in result["message"]
def test_install_base_requirements_failure(api_v3_client):
"""system.py: a pip install that could not start, in the action's output."""
with patch(f"{API}.system._pip_install_requirements", side_effect=OSError(5, LEAK)):
body = _assert_no_leak(api_v3_client.post(
"/api/v3/system/action", json={"action": "install_base_requirements"}))
assert "Failed: OSError:EIO" in body["output"]
+84 -77
View File
@@ -5,20 +5,18 @@ the web UI and the MQTT bridge send) as "restart": with the service running it
ran ``systemctl stop``, slept 1.5s and started it again. Every on-demand or
"Preview on display" click therefore cold-restarted the display process --
every plugin reloaded, the panel blank for seconds -- to deliver a request the
running process takes within a frame over its control socket anyway (see
test_display_pending_changes.py for the display side: commands land
mid-dwell, mid-screen and mid-Vegas-iteration).
running process polls for every ON_DEMAND_POLL_INTERVAL anyway (see
test_on_demand_mailbox.py and test_display_pending_changes.py for the display
side: the mailbox is read mid-dwell, mid-screen and mid-Vegas-iteration).
The restart did not buy anything either: a freshly started display restores
only the on-demand session it saved itself (``display_on_demand_config``).
only the on-demand session it saved itself (``display_on_demand_config``), so
the new request reached it through the same mailbox, one cold start later.
This file previously pinned that restart path (it guarded a broken
``import _pkg.time`` inside it). The path is gone; these tests pin its
replacement: a running service is left alone, a stopped one is started (only
when start_service is set), and the request goes over the control socket
either way -- to a stopped display once it has started and its socket is up,
sent by the web process's dispatcher after the route has answered 202.
Nothing is ever written to the cache: the file mailbox is gone (stage 5).
when start_service is set), and the request lands in the mailbox either way.
The service helpers are patched where they run. display.py binds
_get_display_service_status by value, while _ensure_display_service_running
@@ -27,7 +25,6 @@ _run_systemctl_command is the one place a systemctl command is issued.
"""
import sys
import time
from pathlib import Path
from unittest.mock import patch
@@ -37,8 +34,6 @@ sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
from src.ipc import client as control_client # noqa: E402
START_URL = "/api/v3/display/on-demand/start"
STOP_URL = "/api/v3/display/on-demand/stop"
MAILBOX = "display_on_demand_request"
@@ -51,10 +46,7 @@ def service(api_v3_module):
plugin_manager and config_manager are None so the route skips plugin
resolution (not what is under test here). The cache is the blueprint's
MagicMock cache_manager, so a write would be visible as a set() call.
The control socket answers while the service is active and is missing
(``no_socket``) while it is not; ``sent`` records what it carried.
MagicMock cache_manager, so mailbox writes are visible as set() calls.
"""
api_v3_module.api_v3.plugin_catalog = None
api_v3_module.api_v3.config_manager = None
@@ -70,32 +62,7 @@ def service(api_v3_module):
state["active"] = False
return {"returncode": 0, "stdout": "", "stderr": ""}
sent = []
def socket(action):
def call(request_id, *args, **kwargs):
if not state["active"]:
raise control_client.ControlError("no_socket", "x", sent=False)
sent.append((action, request_id) + args)
return {"accepted": True, "request_id": request_id}
return call
class Clock:
now = 1000.0
def monotonic(self):
return self.now
def time(self):
return self.now
def sleep(self, seconds):
self.now += seconds
with patch("web_interface.blueprints.api_v3.time", Clock()), \
patch(f"{DISPLAY}.control_client.on_demand_start", side_effect=socket("start")), \
patch(f"{DISPLAY}.control_client.on_demand_stop", side_effect=socket("stop")), \
patch("web_interface.blueprints.api_v3._get_display_service_status",
with patch("web_interface.blueprints.api_v3._get_display_service_status",
side_effect=status), \
patch("web_interface.blueprints.api_v3.display._get_display_service_status",
side_effect=status), \
@@ -107,7 +74,6 @@ def service(api_v3_module):
"systemctl": run_systemctl,
"stop_service": stop_service,
"cache": api_v3_module.api_v3.cache_manager,
"sent": sent,
}
@@ -133,15 +99,19 @@ class TestStartWhileTheServiceIsRunning:
assert _systemctl_verbs(service["systemctl"]) == [], (
"a running display service was sent a systemctl command")
def test_the_request_is_sent_to_the_running_display(self, api_v3_client, service):
def test_the_request_is_posted_for_the_running_display(self, api_v3_client, service):
response = api_v3_client.post(
START_URL, json={"plugin_id": "weather", "mode": "weather_current",
"duration": 60, "pinned": True})
data = response.get_json()["data"]
assert service["sent"] == [("start", data["request_id"], "weather",
"weather_current", 60, True)]
assert data["transport"] == "socket"
assert _mailbox_writes(service["cache"]) == []
writes = _mailbox_writes(service["cache"])
assert len(writes) == 1
assert writes[0]["action"] == "start"
assert writes[0]["request_id"] == data["request_id"]
assert writes[0]["plugin_id"] == "weather"
assert writes[0]["mode"] == "weather_current"
assert writes[0]["duration"] == 60
assert writes[0]["pinned"] is True
def test_the_response_reports_the_service_was_not_started(self, api_v3_client, service):
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
@@ -159,20 +129,11 @@ class TestStartWhileTheServiceIsStopped:
def test_start_service_starts_it_once_and_never_stops_it(self, api_v3_client, service):
service["state"]["active"] = False
response = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
# Answered at once: the web process sends the request in the
# background once the started display listens.
assert response.status_code == 202, response.get_json()
assert response.get_json()["status"] == "starting"
assert response.status_code == 200, response.get_json()
assert _systemctl_verbs(service["systemctl"]) == ["start"]
service["stop_service"].assert_not_called()
from web_interface import on_demand_dispatch
dispatcher = on_demand_dispatch.current()
deadline = time.monotonic() + 5
while dispatcher.pending() and time.monotonic() < deadline:
time.sleep(0.01)
assert dispatcher.status()["status"] == "delivered"
assert [s[0] for s in service["sent"]] == ["start"]
assert _mailbox_writes(service["cache"]) == []
# Written before the start, so the new process finds it on its first poll.
assert len(_mailbox_writes(service["cache"])) == 1
def test_without_start_service_it_is_left_stopped(self, api_v3_client, service):
service["state"]["active"] = False
@@ -180,7 +141,6 @@ class TestStartWhileTheServiceIsStopped:
START_URL, json={"plugin_id": "weather", "start_service": "false"})
assert response.status_code == 400
assert _systemctl_verbs(service["systemctl"]) == []
assert service["sent"] == []
def test_a_start_that_fails_is_reported(self, api_v3_client, service):
service["state"]["active"] = False
@@ -191,12 +151,31 @@ class TestStartWhileTheServiceIsStopped:
assert response.get_json()["status"] == "error"
class _Mailbox:
"""The CacheManager calls the routes make, over a dict."""
def __init__(self):
self.entries = {}
def set(self, key, value, ttl=None):
self.entries[key] = value
def get(self, key, max_age=300, memory_ttl=None):
return self.entries.get(key)
def delete(self, key):
self.entries.pop(key, None)
class TestARefusedStartLeavesNoRequestBehind:
"""A start the route answers with an error must not run later.
With the mailbox, the request was posted before the route refused it, and
a display started later ran it. Now nothing is written anywhere: the
request only ever goes over the socket, to a display that answers.
The request was posted (to the mailbox, with the display stopped) before
the route refused it, and the display reads the mailbox for an hour
without looking at a request's age. So "Display service is not running"
(start_service off) or "Failed to start display service" left the
request waiting, and the next time the display started -- minutes later,
by hand -- it ran that plugin, pinned if the request said so.
A socket acknowledgement is the other side of it: the display answered,
so it is running and has the request, whatever systemd says (a display
@@ -205,48 +184,76 @@ class TestARefusedStartLeavesNoRequestBehind:
"""
@pytest.fixture
def stopped(self, service):
def mailbox(self, api_v3_module, service):
box = _Mailbox()
api_v3_module.api_v3.cache_manager = box
service["state"]["active"] = False
return service
return box
@pytest.mark.parametrize("body", [
{"plugin_id": "weather", "start_service": False},
{"plugin_id": "weather"}, # start_service defaults on
])
def test_a_socket_ack_is_a_success_whatever_systemd_says(
self, api_v3_client, stopped, body):
self, api_v3_client, service, mailbox, body):
with patch(f"{DISPLAY}.control_client.on_demand_start",
side_effect=lambda request_id, *a: {"accepted": True}):
response = api_v3_client.post(START_URL, json=body)
assert response.status_code == 200, response.get_json()
assert response.get_json()["data"]["transport"] == "socket"
assert _systemctl_verbs(stopped["systemctl"]) == [], (
assert MAILBOX not in mailbox.entries
assert _systemctl_verbs(service["systemctl"]) == [], (
"a unit was started beside a display that answered the socket")
def test_without_start_service_nothing_is_left_behind(self, api_v3_client, stopped):
def test_without_start_service_the_request_is_taken_back(
self, api_v3_client, service, mailbox):
response = api_v3_client.post(START_URL, json={
"plugin_id": "weather", "pinned": True, "start_service": False})
assert response.status_code == 400
assert response.get_json()["status"] == "error"
assert stopped["cache"].set.call_count == 0
assert stopped["sent"] == []
assert MAILBOX not in mailbox.entries
def test_a_start_that_fails_leaves_nothing_behind(self, api_v3_client, stopped):
stopped["systemctl"].side_effect = lambda args: {
def test_a_start_that_fails_takes_its_request_back(self, api_v3_client, service, mailbox):
service["systemctl"].side_effect = lambda args: {
"returncode": 1, "stdout": "", "stderr": "denied"}
response = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert response.status_code == 500
assert stopped["cache"].set.call_count == 0
assert stopped["sent"] == []
assert MAILBOX not in mailbox.entries
def test_a_newer_request_is_left_alone_on_the_400(self, api_v3_client, service, mailbox):
newer = {"request_id": "someone-else", "action": "start", "plugin_id": "clock"}
def stopped_and_another_post_lands(*args):
mailbox.entries[MAILBOX] = newer
return {"active": False}
with patch(f"{DISPLAY}._get_display_service_status",
side_effect=stopped_and_another_post_lands):
response = api_v3_client.post(START_URL, json={
"plugin_id": "weather", "start_service": False})
assert response.status_code == 400
assert mailbox.entries[MAILBOX] is newer
def test_a_newer_request_is_left_alone_on_the_500(self, api_v3_client, service, mailbox):
newer = {"request_id": "someone-else", "action": "start", "plugin_id": "clock"}
def start_fails_after_another_post(args):
mailbox.entries[MAILBOX] = newer
return {"returncode": 1, "stdout": "", "stderr": "denied"}
service["systemctl"].side_effect = start_fails_after_another_post
response = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert response.status_code == 500
assert mailbox.entries[MAILBOX] is newer
class TestStop:
def test_stop_sends_a_stop_request_and_leaves_the_service_running(
def test_stop_posts_a_stop_request_and_leaves_the_service_running(
self, api_v3_client, service):
response = api_v3_client.post(STOP_URL, json={})
assert response.status_code == 200, response.get_json()
assert [s[0] for s in service["sent"]] == ["stop"]
assert _mailbox_writes(service["cache"]) == []
writes = _mailbox_writes(service["cache"])
assert [w["action"] for w in writes] == ["stop"]
service["stop_service"].assert_not_called()
assert _systemctl_verbs(service["systemctl"]) == []
+83 -396
View File
@@ -1,12 +1,15 @@
"""POST /display/on-demand/start and /stop: the control socket is the only way.
"""POST /display/on-demand/start and /stop: control socket first, mailbox fallback.
The routes hand the request to the display over the control socket
(src/ipc) and get an acknowledgement. The file mailbox they used to fall
back to (``display_on_demand_request``) is gone (stage 5), so nothing is
ever written to the cache. When no display is listening (a stopped one, or
one still starting) the start route starts the service if asked to and
sends the request again once the socket is up; every other failure is
answered as what it is.
(src/ipc) and get an acknowledgement. When the socket could not carry the
request -- no socket (a stopped display, or one older than the socket), a
refused or timed-out connect, a display too old to know the command, a bug
in the client -- they write the file mailbox exactly as they did before the
socket existed. When the display had the request and failed it (a full
queue, bad arguments, no answer in time) the route says so and writes
nothing. These tests pin each path, that at most one of them is used, that
the response says which, and that the request id is the same either way (the
display deduplicates on it).
The socket client is patched at the route's module attribute; the last class
runs a real server on a temp socket (Linux/macOS only).
@@ -14,8 +17,6 @@ runs a real server on a temp socket (Linux/macOS only).
import os
import sys
import time
import types
from pathlib import Path
from unittest.mock import patch
@@ -32,12 +33,11 @@ START_URL = "/api/v3/display/on-demand/start"
STOP_URL = "/api/v3/display/on-demand/stop"
MAILBOX = "display_on_demand_request"
CLIENT = "web_interface.blueprints.api_v3.display.control_client"
DISPLAY = "web_interface.blueprints.api_v3.display"
@pytest.fixture
def service(api_v3_module):
"""A running display service; records systemctl calls and cache writes."""
"""A running display service; records systemctl calls and mailbox writes."""
api_v3_module.api_v3.plugin_catalog = None
api_v3_module.api_v3.config_manager = None
state = {"active": True}
@@ -101,398 +101,85 @@ class TestSocketPath:
assert _mailbox_writes(service["cache"]) == []
class FakeTime:
"""Stands in for the route's ``time`` module: sleep() moves the clock."""
def __init__(self):
self.now = 1000.0
self.sleeps = 0
def monotonic(self):
return self.now
def time(self):
return self.now
def sleep(self, seconds):
self.sleeps += 1
self.now += seconds
@pytest.fixture
def clock():
fake = FakeTime()
with patch("web_interface.blueprints.api_v3.time", fake):
yield fake
def _attempts(*outcomes, calls=None, clock=None):
"""A fake on_demand_start/stop that answers ``outcomes`` in turn (an
exception is raised, anything else acks); the last one repeats."""
outcomes = list(outcomes)
seen = calls if calls is not None else []
def attempt(request_id, *a, **kw):
seen.append(clock.now if clock is not None else None)
outcome = outcomes.pop(0) if len(outcomes) > 1 else outcomes[0]
if isinstance(outcome, BaseException):
raise outcome
return _ack(request_id)
return attempt
def _until(predicate, timeout=5.0):
end = time.monotonic() + timeout
while time.monotonic() < end:
if predicate():
return True
time.sleep(0.005)
return False
def _no_socket():
return control_client.ControlError("no_socket", "x", sent=False)
class TestNoDisplayListening:
"""No socket to talk to: the display is stopped, or still starting.
The route answers at once (202, ``status: "starting"``) and the web
process's dispatcher (web_interface/on_demand_dispatch.py) sends the
request until the display acknowledges it; the status routes report the
outcome. The dispatcher runs on real time with short waits here.
"""
@pytest.fixture(autouse=True)
def quick(self, monkeypatch, service):
from web_interface import on_demand_dispatch
monkeypatch.setattr(on_demand_dispatch, "RETRY_INTERVAL", 0.01)
monkeypatch.setattr(on_demand_dispatch, "START_WAIT_SECONDS", 2.0)
monkeypatch.setattr(f"{DISPLAY}.ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS", 1.0)
service["cache"].get.return_value = None # the display has published nothing
@staticmethod
def _outcome(status=None):
from web_interface import on_demand_dispatch
d = on_demand_dispatch.current()
assert d is not None
assert _until(lambda: not d.pending()), "the dispatcher never finished"
return d.status()
def test_a_display_still_starting_gets_the_request_once_it_listens(
self, api_v3_client, service):
calls = []
with patch(f"{CLIENT}.on_demand_start",
side_effect=_attempts(_no_socket(), _no_socket(), "ack", calls=calls)):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 202, resp.get_json()
body = resp.get_json()
assert body["status"] == "starting"
data = body["data"]
assert data["pending"] is True and data["socket_error"] == "no_socket"
assert data["service"]["started"] is False
outcome = self._outcome()
assert outcome["status"] == "delivered"
assert outcome["request_id"] == data["request_id"]
assert len(calls) == 3
assert not [c for c in service["calls"] if c[0] == "systemctl"]
assert _mailbox_writes(service["cache"]) == []
def test_a_stopped_display_is_started_and_answered_at_once(self, api_v3_client, service):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_start",
side_effect=_attempts(*[_no_socket()] * 6, "ack")) as start:
started = time.monotonic()
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather",
"duration": 30})
answered = time.monotonic() - started
assert resp.status_code == 202, resp.get_json()
data = resp.get_json()["data"]
assert data["service"]["started"] is True
from web_interface import on_demand_dispatch
assert data["wait_seconds"] == on_demand_dispatch.START_WAIT_SECONDS
outcome = self._outcome()
assert answered < 1.0, "the route waited for the display"
assert service["calls"] == [("systemctl", "start")]
assert outcome["status"] == "delivered"
# The same request, the same id, every time.
assert {call.args[0] for call in start.call_args_list} == {data["request_id"]}
assert _mailbox_writes(service["cache"]) == []
def test_while_pending_the_status_routes_say_starting(self, api_v3_client, service):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
rid = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) \
.get_json()["data"]["request_id"]
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
assert status["source"] == "web"
assert status["state"]["status"] == "starting"
assert status["state"]["request_id"] == rid
assert status["state"]["plugin_id"] == "weather"
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
assert current["on_demand_pending"]["status"] == "starting"
def test_a_display_that_never_comes_up_is_a_start_timeout(self, api_v3_client, service):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 202
outcome = self._outcome()
assert outcome["status"] == "error" and outcome["error"] == "start-timeout"
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
assert status["state"]["status"] == "error"
assert status["state"]["error"] == "start-timeout"
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
assert current["on_demand_pending"]["error"] == "start-timeout"
def test_a_later_display_state_replaces_the_failure(self, api_v3_client, service):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
api_v3_client.post(START_URL, json={"plugin_id": "weather"})
failed_at = self._outcome()["last_updated"]
later = {"active": True, "status": "active", "plugin_id": "clock",
"last_updated": failed_at + 5}
service["cache"].get.return_value = later
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
assert status["state"]["plugin_id"] == "clock" and status["source"] == "cache"
def test_a_running_service_without_a_socket_is_waited_for_less(
self, api_v3_client, service):
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 202
assert resp.get_json()["data"]["wait_seconds"] == 1.0
assert self._outcome()["error"] == "start-timeout"
assert not [c for c in service["calls"] if c[0] == "systemctl"]
def test_a_different_failure_while_waiting_ends_it(self, api_v3_client, service):
calls = []
busy = control_client.ControlError("busy", "x", sent=True)
with patch(f"{CLIENT}.on_demand_start",
side_effect=_attempts(_no_socket(), _no_socket(), busy, calls=calls)):
assert api_v3_client.post(START_URL, json={"plugin_id": "weather"}).status_code == 202
outcome = self._outcome()
assert outcome["status"] == "error" and outcome["error"] == "busy"
assert len(calls) == 3
def test_a_stop_while_pending_cancels_it(self, api_v3_client, service):
service["state"]["active"] = False
calls = []
with patch(f"{CLIENT}.on_demand_start",
side_effect=_attempts(_no_socket(), calls=calls)), \
patch(f"{CLIENT}.on_demand_stop", side_effect=_no_socket()):
rid = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) \
.get_json()["data"]["request_id"]
resp = api_v3_client.post(STOP_URL, json={})
assert resp.status_code == 200, resp.get_json()
data = resp.get_json()["data"]
assert data["cancelled_request_id"] == rid
outcome = self._outcome()
n = len(calls)
time.sleep(0.05)
assert len(calls) == n, "the cancelled start was still being sent"
assert outcome["status"] == "idle" and outcome["last_event"] == "requested-stop"
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
assert status["state"]["status"] == "idle" and status["source"] != "web"
# -- delivered, but the display has not acted on it yet ----------------
# The display acknowledges a start when its socket opens and acts on it
# seconds later (Vegas builds its first strip first: ~5 s on ledpi);
# meanwhile it publishes its own idle state, which must not flash.
def _delivered(self, api_v3_client, service):
with patch(f"{CLIENT}.on_demand_start",
side_effect=_attempts(_no_socket(), "ack")):
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) .get_json()["data"]
outcome = self._outcome()
assert outcome["status"] == "delivered"
return data["request_id"], outcome["last_updated"]
@staticmethod
def _status(api_v3_client):
return api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
def test_a_delivered_start_stays_starting_until_the_display_answers(
self, api_v3_client, service):
rid, delivered_at = self._delivered(api_v3_client, service)
# The display's startup state: idle, published after the ack, for
# no request (or an older one).
for older in (None, "an-older-request"):
service["cache"].get.return_value = {
"active": False, "status": "idle", "request_id": older,
"last_updated": delivered_at + 1}
data = self._status(api_v3_client)
assert data["source"] == "web", older
assert data["state"]["status"] == "starting"
assert data["state"]["delivered"] is True
assert data["state"]["request_id"] == rid
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
assert current["on_demand_pending"]["delivered"] is True
def test_the_display_state_for_the_request_takes_over(self, api_v3_client, service):
rid, delivered_at = self._delivered(api_v3_client, service)
service["cache"].get.return_value = {
"active": True, "status": "active", "plugin_id": "weather",
"request_id": rid, "last_updated": delivered_at + 5}
data = self._status(api_v3_client)
assert data["source"] == "cache" and data["state"]["status"] == "active"
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
assert "on_demand_pending" not in current
def test_its_error_takes_over_too(self, api_v3_client, service):
rid, delivered_at = self._delivered(api_v3_client, service)
service["cache"].get.return_value = {
"active": False, "status": "error", "error": "load-failed",
"request_id": rid, "last_updated": delivered_at + 5}
assert self._status(api_v3_client)["state"]["error"] == "load-failed"
def test_an_older_display_answers_with_any_newer_state(self, api_v3_client, service):
# A display before request_id was published: timestamps decide.
_, delivered_at = self._delivered(api_v3_client, service)
service["cache"].get.return_value = {"active": False, "status": "idle",
"last_updated": delivered_at - 1}
assert self._status(api_v3_client)["source"] == "web"
service["cache"].get.return_value = {"active": True, "status": "active",
"last_updated": delivered_at + 1}
assert self._status(api_v3_client)["source"] == "cache"
def test_a_delivered_start_is_shown_for_a_limited_time(
self, api_v3_client, service, monkeypatch):
from web_interface import on_demand_dispatch
_, delivered_at = self._delivered(api_v3_client, service)
service["cache"].get.return_value = {"active": False, "status": "idle",
"request_id": None,
"last_updated": delivered_at + 1}
cap = on_demand_dispatch.DELIVERED_SHOWN_SECONDS
assert cap == 30.0
clock = types.SimpleNamespace(time=lambda: delivered_at + cap - 1,
monotonic=time.monotonic, sleep=time.sleep)
monkeypatch.setattr("web_interface.blueprints.api_v3.time", clock)
assert self._status(api_v3_client)["source"] == "web"
clock.time = lambda: delivered_at + cap + 1
data = self._status(api_v3_client)
assert data["source"] == "cache" and data["state"]["status"] == "idle"
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
assert "on_demand_pending" not in current
def test_a_new_start_supersedes_the_pending_one(self, api_v3_client, service):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
old = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) \
.get_json()["data"]["request_id"]
sent = []
service["state"]["active"] = True
with patch(f"{CLIENT}.on_demand_start",
side_effect=lambda rid, *a: sent.append(rid) or {"accepted": True}):
resp = api_v3_client.post(START_URL, json={"plugin_id": "clock"})
assert resp.status_code == 200
outcome = self._outcome()
# The old start is cancelled before the new one is sent, and
# cancel() waits out a send of it in flight: nothing of it lands
# after the new one.
new = resp.get_json()["data"]["request_id"]
assert sent[-1] == new and sent.count(new) == 1
assert outcome["status"] == "idle" and outcome["last_event"] == "superseded"
def test_a_service_that_will_not_start_is_an_error(self, api_v3_client, service, clock):
service["state"]["active"] = False
with patch("web_interface.blueprints.api_v3._run_systemctl_command",
return_value={"returncode": 1, "stdout": "", "stderr": "nope"}), \
patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()) as start:
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 500
assert "Failed to start display service" in resp.get_json()["message"]
assert start.call_count == 1
def test_stop_with_no_display_running_is_an_error(self, api_v3_client, service, clock):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_stop", side_effect=_no_socket()) as stop:
resp = api_v3_client.post(STOP_URL, json={})
assert resp.status_code == 503
body = resp.get_json()
assert "not running" in body["message"]
assert body["data"]["socket_error"] == "no_socket"
assert stop.call_count == 1 and clock.sleeps == 0
assert service["calls"] == []
assert _mailbox_writes(service["cache"]) == []
def test_stop_with_a_running_service_but_no_socket_is_an_error(
self, api_v3_client, service, clock):
with patch(f"{CLIENT}.on_demand_stop",
side_effect=control_client.ControlError("refused", "x")):
resp = api_v3_client.post(STOP_URL, json={})
assert resp.status_code == 503
assert "still be starting" in resp.get_json()["message"]
def test_stop_with_stop_service_stops_it_anyway(self, api_v3_client, service, clock):
with patch(f"{CLIENT}.on_demand_stop", side_effect=_no_socket()), \
patch(f"{DISPLAY}._stop_display_service",
return_value={"active": False}) as stop:
resp = api_v3_client.post(STOP_URL, json={"stop_service": True})
assert resp.status_code == 200
assert resp.get_json()["data"]["socket_error"] == "no_socket"
stop.assert_called_once()
assert _mailbox_writes(service["cache"]) == []
def test_the_socket_is_off_in_the_test_suite(self, api_v3_client, service):
# conftest's _hermetic_control_socket: a suite run on a device must
# not drive the live display. Nothing is retried or started for it.
assert os.environ[c.SOCKET_PATH_ENV] == "off"
service["state"]["active"] = False
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 503
assert resp.get_json()["data"]["socket_error"] in ("disabled", "unsupported")
assert service["calls"] == []
class TestNothingElseIsRetried:
@pytest.mark.parametrize("reason,sent,status", [
("timeout", False, 503), ("busy", False, 503), ("forbidden", False, 503),
("invalid_request", False, 503), ("disabled", False, 503),
("unsupported", False, 503), ("unknown_command", True, 503),
("unsupported_version", True, 503),
class TestMailboxFallback:
@pytest.mark.parametrize("reason", [
"no_socket", "refused", "timeout", "closed", "bad_response", "invalid_request",
"busy", "forbidden", "unknown_command", "unsupported_version", "disabled",
"unsupported",
])
def test_answered_at_once_and_nothing_written(self, api_v3_client, service, clock,
reason, sent, status):
service["state"]["active"] = False
def test_a_request_the_socket_never_carried_writes_the_mailbox(
self, api_v3_client, service, reason):
# sent=False: the display never had it (no socket, a refused or
# timed-out connect, turned away at the door).
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError(reason, "x", sent=sent)) as start:
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == status
body = resp.get_json()
assert body["status"] == "error"
assert body["data"]["socket_error"] == reason
assert start.call_count == 1 and clock.sleeps == 0
assert service["calls"] == []
side_effect=control_client.ControlError(reason, "x", sent=False)):
resp = api_v3_client.post(START_URL, json={
"plugin_id": "weather", "mode": "weather_current",
"duration": 60, "pinned": True})
assert resp.status_code == 200
data = resp.get_json()["data"]
assert data["transport"] == "mailbox"
assert data["socket_error"] == reason
[write] = _mailbox_writes(service["cache"])
assert write["request_id"] == data["request_id"]
assert write["action"] == "start"
assert (write["plugin_id"], write["mode"], write["duration"], write["pinned"]) == \
("weather", "weather_current", 60, True)
def test_a_client_bug_is_an_error(self, api_v3_client, service):
def test_a_client_bug_still_falls_back(self, api_v3_client, service):
with patch(f"{CLIENT}.on_demand_start", side_effect=RuntimeError("boom")):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 503
assert resp.status_code == 200
assert resp.get_json()["data"]["socket_error"] == "internal"
assert _mailbox_writes(service["cache"]) == []
assert len(_mailbox_writes(service["cache"])) == 1
def test_an_unknown_reason_is_reported_as_other(self, api_v3_client, service):
# Only known codes are echoed back; anything else stays server-side.
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError("/run/secret/path", "x")):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 503
assert resp.get_json()["data"]["socket_error"] == "other"
assert "/run/secret" not in resp.get_data(as_text=True)
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
assert data["transport"] == "mailbox"
assert data["socket_error"] == "other"
assert len(_mailbox_writes(service["cache"])) == 1
def test_every_display_error_code_is_reportable(self):
from web_interface.blueprints.api_v3 import display
codes = {v for k, v in vars(c.ErrorCode).items() if not k.startswith("_")}
assert codes <= set(display._REPORTABLE_SOCKET_REASONS)
def test_stop_falls_back(self, api_v3_client, service):
with patch(f"{CLIENT}.on_demand_stop",
side_effect=control_client.ControlError("timeout")):
data = api_v3_client.post(STOP_URL, json={}).get_json()["data"]
assert data["transport"] == "mailbox"
[write] = _mailbox_writes(service["cache"])
assert write == {"request_id": data["request_id"], "action": "stop",
"timestamp": write["timestamp"]}
def test_a_stopped_display_gets_the_mailbox_before_it_is_started(
self, api_v3_client, service):
service["state"]["active"] = False
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError("no_socket")):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == 200
assert service["calls"] == [("cache", MAILBOX), ("systemctl", "start")]
def test_the_socket_is_off_in_the_test_suite(self, api_v3_client, service):
# conftest's _hermetic_control_socket: a suite run on a device must
# not drive the live display.
assert os.environ[c.SOCKET_PATH_ENV] == "off"
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
assert data["transport"] == "mailbox"
assert data["socket_error"] in ("disabled", "unsupported") # Linux, Windows
class TestTheDisplayHadIt:
"""Once the display has the request, its answer stands.
"""Once the display has the request, its answer stands: no mailbox copy.
A busy queue, a refusal or silence after the request was sent mean the
display may have applied it, so the route reports the failure and does
not send it again.
display may have applied it, or would refuse the mailbox copy too, so
the route reports the failure instead of posting it a second time.
"""
@pytest.mark.parametrize("reason,status", [
@@ -530,6 +217,16 @@ class TestTheDisplayHadIt:
stop.assert_called_once()
assert _mailbox_writes(service["cache"]) == []
@pytest.mark.parametrize("reason", ["unknown_command", "unsupported_version"])
def test_an_older_display_that_does_not_speak_it_gets_the_mailbox(
self, api_v3_client, service, reason):
# The upgrade case: new web interface, display still on an old build.
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError(reason, "x", sent=True)):
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
assert data["transport"] == "mailbox" and data["socket_error"] == reason
assert len(_mailbox_writes(service["cache"])) == 1
@pytest.mark.skipif(not c.socket_supported(), reason="AF_UNIX sockets are Linux/macOS only")
class TestRealSocket:
@@ -562,21 +259,11 @@ class TestRealSocket:
assert data["transport"] == "socket"
assert [x.request_id for x in live.drain()] == [data["request_id"]]
def test_a_display_that_went_away_is_given_up_on(self, api_v3_client, service, live,
monkeypatch):
from web_interface import on_demand_dispatch
monkeypatch.setattr(f"{DISPLAY}.ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS", 0.3)
monkeypatch.setattr(on_demand_dispatch, "RETRY_INTERVAL", 0.01)
def test_a_display_that_went_away_falls_back(self, api_v3_client, service, live):
live.close()
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
# The service still reads as running: answered at once, and the
# web process gives up once the wait is over.
assert resp.status_code == 202
assert resp.get_json()["data"]["socket_error"] == "no_socket"
d = on_demand_dispatch.current()
assert _until(lambda: not d.pending())
assert d.status()["error"] == "start-timeout"
assert _mailbox_writes(service["cache"]) == []
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
assert data["transport"] == "mailbox" and data["socket_error"] == "no_socket"
assert len(_mailbox_writes(service["cache"])) == 1
def test_a_full_queue_is_reported_not_mailed(self, api_v3_client, service, monkeypatch):
import shutil
+4 -3
View File
@@ -95,9 +95,10 @@ class TestRefreshPluginStore:
RuntimeError("failed at /home/user/LEDMatrix/src/secret.py line 42"))
body = api_v3_client.post(self.URL, json={}).get_json()
assert "Traceback" not in str(body)
# `details` is describe_exception output: one line, type-named,
# credential-redacted. It may quote the message, but never a stack.
assert body["details"].startswith("RuntimeError:")
# `details` is describe_exception output: the type, never the
# message or a stack.
assert body["details"] == "RuntimeError"
assert "secret.py" not in str(body)
assert "\n" not in body["details"]
+37 -36
View File
@@ -6,7 +6,7 @@ a Vegas iteration runs for vegas_scroll.max_cycle_duration (240s here). On a
real Pi on 2026-09-23:
* an on-demand request posted at 10:54:27 was activated at 10:57:24, when the
Vegas iteration it arrived during finally ended -- nothing read the request
Vegas iteration it arrived during finally ended -- nothing read the mailbox
in between, because _check_vegas_interrupt only looked at a flag that the
main-loop read sets;
* two brightness saves 12s apart inside one 30s screen never showed at all.
@@ -15,11 +15,6 @@ _service_pending_changes is the fix: a throttled pass the dwell sleep, the
render loops and the Vegas interrupt check all call. These tests drive those
long stretches on a fake clock and check a change lands within one throttle
interval -- and that the throttle holds, since the callers run at frame rate.
The on-demand requests here come from a plugin in the display process
(submit_plugin_on_demand), the way in that needs no socket; the file mailbox
these tests used to write is gone (stage 5). A queued request skips the
throttle, so it lands at the next pass's call, not the next interval.
"""
import threading
@@ -33,6 +28,7 @@ from src.vegas_mode.config import VegasModeConfig
from src.vegas_mode.coordinator import VegasModeCoordinator
VEGAS_ITERATION_SECONDS = 240
REQUEST_KEY = 'display_on_demand_request'
class FakeClock:
@@ -79,15 +75,26 @@ def clock(monkeypatch):
@pytest.fixture
def controller(test_display_controller, clock):
"""A controller at rest: no schedule, full brightness, nothing queued."""
"""A controller at rest: no schedule, full brightness, empty mailbox."""
c = test_display_controller
c._refresh_config_cache({'display': {'hardware': {'brightness': 90}}})
c.current_brightness = 90
c.is_display_active = True
c._check_wifi_status_message = MagicMock(return_value=None)
c.cache_manager.get = MagicMock(return_value=None)
c.cache_manager.delete = MagicMock()
c.mailbox = {} # what the web process has written
def cache_get(key, *args, **kwargs):
if key == REQUEST_KEY:
return c.mailbox.get('request')
return None
def cache_delete(key):
if key == REQUEST_KEY:
c.mailbox.pop('request', None)
c.cache_manager.get = MagicMock(side_effect=cache_get)
c.cache_manager.delete = MagicMock(side_effect=cache_delete)
c.cache_manager.set = MagicMock()
c.display_manager.set_brightness = MagicMock(return_value=True)
c.display_manager.update_display = MagicMock()
@@ -103,12 +110,11 @@ def controller(test_display_controller, clock):
return c
def post_request(controller, request_id='r1', mode='clock', action='start'):
"""A plugin asking for the screen (or giving it back) from its thread."""
assert controller.submit_plugin_on_demand({
'request_id': request_id, 'action': action,
'plugin_id': mode, 'mode': mode, 'source': 'plugin',
})
def post_request(controller, request_id='r1', mode='clock'):
controller.mailbox['request'] = {
'request_id': request_id, 'action': 'start',
'plugin_id': mode, 'mode': mode,
}
def save_brightness(controller, brightness):
@@ -118,11 +124,9 @@ def save_brightness(controller, brightness):
{'display': {'hardware': {'brightness': brightness}}})
def count_passes(controller):
"""Count the pending-changes passes that got past the throttle."""
controller._poll_on_demand_requests = MagicMock(
wraps=controller._poll_on_demand_requests)
return controller._poll_on_demand_requests
def mailbox_reads(controller):
return sum(1 for call in controller.cache_manager.get.call_args_list
if call.args and call.args[0] == REQUEST_KEY)
def vegas_coordinator(controller):
@@ -203,16 +207,15 @@ class TestOnDemandDuringVegas:
"interrupt checker never saw the request")
assert controller.on_demand_active
def test_a_quiet_iteration_runs_out(self, controller, clock):
def test_an_empty_mailbox_lets_the_iteration_run_out(self, controller, clock):
coord = vegas_coordinator(controller)
passes = count_passes(controller)
assert coord.run_iteration() is True
assert not controller.on_demand_active
# And the pass is throttled: at most one per interval, not per check.
max_passes = VEGAS_ITERATION_SECONDS / controller.PENDING_CHANGES_INTERVAL + 1
assert 0 < passes.call_count <= max_passes
# And the read is throttled: at most one per interval, not per check.
max_reads = VEGAS_ITERATION_SECONDS / controller.PENDING_CHANGES_INTERVAL + 1
assert 0 < mailbox_reads(controller) <= max_reads
checks = coord.frames // coord._interrupt_check_interval
assert passes.call_count < checks / 2
assert mailbox_reads(controller) < checks / 2
class TestBrightnessIsAppliedMidScreen:
@@ -292,25 +295,23 @@ class TestBrightnessIsAppliedMidScreen:
class TestThrottle:
def test_no_passes_or_brightness_calls_between_passes(self, controller, clock):
def test_no_reads_or_brightness_calls_between_passes(self, controller, clock):
controller._service_pending_changes()
passes = count_passes(controller)
reads = mailbox_reads(controller)
save_brightness(controller, 40)
post_request(controller)
for _ in range(500): # a few seconds of frames, all inside one interval
controller._service_pending_changes()
clock.t += controller.PENDING_CHANGES_INTERVAL / 1000
assert passes.call_count == 0
assert mailbox_reads(controller) == reads
controller.display_manager.set_brightness.assert_not_called()
assert not controller.on_demand_active
clock.t += controller.PENDING_CHANGES_INTERVAL
controller._service_pending_changes()
assert passes.call_count == 1
# The poll, plus the consume step's re-read of the request it acted on.
assert mailbox_reads(controller) > reads
controller.display_manager.set_brightness.assert_called_once_with(40)
def test_a_queued_request_skips_the_throttle(self, controller, clock):
controller._service_pending_changes()
post_request(controller)
controller._service_pending_changes() # inside the interval
assert controller.on_demand_active
def test_between_passes_it_does_no_work_at_all(self, controller, clock):
@@ -397,7 +398,7 @@ class TestScheduleAndDwells:
controller._service_pending_changes()
assert controller.on_demand_active
clock.t += controller.PENDING_CHANGES_INTERVAL
post_request(controller, request_id='r2', action='stop')
controller.mailbox['request'] = {'request_id': 'r2', 'action': 'stop'}
start = clock.t
controller._sleep_with_plugin_updates(30)
assert not controller.on_demand_active
+203 -130
View File
@@ -24,7 +24,7 @@ sys.path.insert(0, str(Path(__file__).parent.parent))
from src.cache_manager import CacheManager # noqa: E402
from src import error_aggregator as errors # noqa: E402
from src.error_aggregator import ( # noqa: E402
ERROR_SNAPSHOT_KEY, ErrorAggregator,
ERROR_CLEAR_REQUEST_KEY, ERROR_SNAPSHOT_KEY, ErrorAggregator,
ErrorSnapshotPublisher,
)
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
@@ -314,36 +314,151 @@ class TestRoutes:
assert secret not in published, secret
class TestClear:
def test_clear_is_applied_by_the_display_and_republished(self, web, display):
aggregator, publisher, _ = display
_fail(aggregator)
_fail(aggregator)
publisher.tick()
response = web.post("/api/v3/errors/clear", json={"all": True})
assert response.status_code == 200
body = response.get_json()["data"]
assert body["clear_requested"] is True
assert body["cleared_count"] == 2
# The display applies it on its next tick, throttle or not...
assert publisher.tick() is True
assert aggregator.get_error_summary()["total_errors"] == 0
data = _summary(web)
assert data["total_errors"] == 0 and data["clear_pending"] is False
# ...and only once.
assert publisher.tick() is False
def test_summary_hides_cleared_errors_before_the_display_applies_it(self, web, display):
aggregator, publisher, _ = display
for _ in range(5):
_fail(aggregator)
publisher.tick()
web.post("/api/v3/errors/clear", json={"all": True})
# No display tick yet.
data = _summary(web)
assert data["clear_pending"] is True
assert data["total_errors"] == 0
assert data["recent_errors"] == [] and data["active_patterns"] == {}
assert data["plugin_error_counts"] == {}
plugin = web.get("/api/v3/errors/plugin/p1").get_json()["data"]
assert plugin["status"] == "healthy" and plugin["total_errors"] == 0
def test_a_snapshot_written_just_before_the_clear_cannot_bring_errors_back(
self, web, display, shared_cache):
# The race: the display builds a snapshot, the user clicks Clear, and
# the display's write lands after the request.
aggregator, publisher, _ = display
display_cache, _, _ = shared_cache
for _ in range(3):
_fail(aggregator)
stale = aggregator.build_snapshot()
web.post("/api/v3/errors/clear", json={"all": True})
stale["applied_clear_id"] = None
display_cache.set(ERROR_SNAPSHOT_KEY, stale)
assert _summary(web)["total_errors"] == 0
def test_errors_after_the_clear_are_kept(self, web, display):
aggregator, publisher, clock = display
_fail(aggregator, plugin_id="before")
publisher.tick()
web.post("/api/v3/errors/clear", json={"all": True})
# An error lands after the request but before the display applies it.
for record in aggregator._records:
record.timestamp -= timedelta(seconds=5)
_fail(aggregator, plugin_id="after")
publisher.tick()
data = _summary(web)
assert data["plugin_error_counts"] == {"after": {"ValueError": 1}}
assert data["clear_pending"] is False
def test_age_based_clear(self, web, display):
aggregator, publisher, _ = display
_fail(aggregator, plugin_id="old")
_fail(aggregator, plugin_id="old")
for record in aggregator._records:
record.timestamp -= timedelta(hours=3)
_fail(aggregator, plugin_id="new")
publisher.tick()
body = web.post("/api/v3/errors/clear", json={"max_age_hours": 1}).get_json()["data"]
assert body["cleared_count"] == 2
pending = _summary(web)
assert pending["clear_pending"] is True
assert [r["plugin_id"] for r in pending["recent_errors"]] == ["new"]
publisher.tick()
data = _summary(web)
assert data["plugin_error_counts"] == {"new": {"ValueError": 1}}
assert data["total_errors"] == 1
def test_a_narrower_clear_does_not_undo_a_pending_wider_one(self, web, display):
aggregator, publisher, _ = display
for _ in range(3):
_fail(aggregator)
publisher.tick()
web.post("/api/v3/errors/clear", json={"all": True})
web.post("/api/v3/errors/clear", json={"max_age_hours": 24})
assert _summary(web)["total_errors"] == 0
publisher.tick()
assert aggregator.get_error_summary()["total_errors"] == 0
def test_default_body_still_means_older_than_24_hours(self, web, display):
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
response = web.post("/api/v3/errors/clear")
assert response.status_code == 200
assert response.get_json()["data"]["cleared_count"] == 0
publisher.tick()
assert _summary(web)["total_errors"] == 1
def test_validation_is_unchanged_but_all_skips_it(self, web):
assert web.post("/api/v3/errors/clear", json={"max_age_hours": 0}).status_code == 400
assert web.post("/api/v3/errors/clear", json={"max_age_hours": 9000}).status_code == 400
assert web.post("/api/v3/errors/clear",
json={"all": True, "max_age_hours": "junk"}).status_code == 200
def test_a_request_that_did_not_reach_the_cache_is_an_error(self, web, api_v3_module): # noqa: F811
cache = MagicMock()
cache.get.return_value = None
api_v3_module.api_v3.cache_manager = cache
response = web.post("/api/v3/errors/clear", json={"all": True})
assert response.status_code == 500
assert "clear request" in response.get_json()["message"]
CLIENT = "web_interface.blueprints.api_v3.control_client"
RETIRED_CLEAR_KEY = "plugin_error_clear_request"
def _mailbox_file(shared_cache):
_, _, directory = shared_cache
return directory / f"{RETIRED_CLEAR_KEY}.json"
return directory / f"{ERROR_CLEAR_REQUEST_KEY}.json"
@pytest.fixture
def socket_up(display, monkeypatch):
"""The control socket, as the display serves it: errors_clear runs the
display's own handler against its publisher."""
from src.ipc import client as control_client
from src.ipc.contract import ErrorsClearArgs
_, publisher, _ = display
monkeypatch.setattr(errors, "_snapshot_publisher", publisher)
calls = []
class TestClearOverTheSocket:
"""``errors.clear``: the display applies the clear before it answers, and
the mailbox is written only when the socket could not carry it."""
def errors_clear(request_id, cutoff, **kw):
calls.append((request_id, cutoff))
return errors.apply_error_clear(request_id, ErrorsClearArgs(cutoff=cutoff))
@pytest.fixture
def socket_up(self, display, monkeypatch):
"""The control socket, as the display serves it: errors_clear runs
the display's own handler against its publisher."""
from src.ipc import client as control_client
from src.ipc.contract import ErrorsClearArgs
_, publisher, _ = display
monkeypatch.setattr(errors, "_snapshot_publisher", publisher)
calls = []
monkeypatch.setattr(f"{CLIENT}.errors_clear", errors_clear)
assert control_client.errors_clear is errors_clear
return calls
def errors_clear(request_id, cutoff, **kw):
calls.append((request_id, cutoff))
return errors.apply_error_clear(request_id, ErrorsClearArgs(cutoff=cutoff))
class TestClear:
"""``errors.clear``: the display applies the clear before it answers."""
monkeypatch.setattr(f"{CLIENT}.errors_clear", errors_clear)
assert control_client.errors_clear is errors_clear
return calls
def test_socket_clear_is_applied_before_the_answer(self, web, display, socket_up,
shared_cache):
@@ -356,7 +471,6 @@ class TestClear:
body = response.get_json()
data = body["data"]
assert data["transport"] == "socket" and data["applied"] is True
assert data["clear_requested"] is True
assert data["cleared_count"] == 3
assert body["message"] == "Cleared all errors"
[(request_id, _)] = socket_up
@@ -365,134 +479,90 @@ class TestClear:
assert aggregator.get_error_summary()["total_errors"] == 0
summary = _summary(web)
assert summary["total_errors"] == 0 and summary["clear_pending"] is False
# And no mailbox file.
assert not _mailbox_file(shared_cache).exists()
# Nothing left for a tick to do.
assert publisher.tick() is False
def test_errors_after_the_clear_are_kept(self, web, display, socket_up):
aggregator, publisher, _ = display
_fail(aggregator, plugin_id="before")
for record in aggregator._records:
record.timestamp -= timedelta(seconds=5)
publisher.tick()
web.post("/api/v3/errors/clear", json={"all": True})
_fail(aggregator, plugin_id="after")
publisher.min_interval = 0
publisher.tick()
data = _summary(web)
assert data["plugin_error_counts"] == {"after": {"ValueError": 1}}
assert data["clear_pending"] is False
def test_age_based_clear(self, web, display, socket_up):
aggregator, publisher, _ = display
_fail(aggregator, plugin_id="old")
_fail(aggregator, plugin_id="old")
for record in aggregator._records:
record.timestamp -= timedelta(hours=3)
_fail(aggregator, plugin_id="new")
publisher.tick()
body = web.post("/api/v3/errors/clear", json={"max_age_hours": 1}).get_json()["data"]
assert body["cleared_count"] == 2
data = _summary(web)
assert data["plugin_error_counts"] == {"new": {"ValueError": 1}}
assert data["total_errors"] == 1
def test_default_body_still_means_older_than_24_hours(self, web, display, socket_up):
def test_an_older_mailbox_request_does_not_read_as_pending(self, web, display, socket_up,
shared_cache, monkeypatch):
# A clear that went to the mailbox while the socket was down, then a
# wider one over the socket: the old request has nothing left to hide.
from unittest.mock import patch
from src.ipc import client as control_client
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
response = web.post("/api/v3/errors/clear")
assert response.status_code == 200
assert response.get_json()["data"]["cleared_count"] == 0
assert _summary(web)["total_errors"] == 1
with patch(f"{CLIENT}.errors_clear",
side_effect=control_client.ControlError("no_socket")):
data = web.post("/api/v3/errors/clear", json={"max_age_hours": 1}).get_json()["data"]
assert data["transport"] == "mailbox"
assert _mailbox_file(shared_cache).exists()
data = web.post("/api/v3/errors/clear", json={"all": True}).get_json()["data"]
assert data["transport"] == "socket"
assert _summary(web)["clear_pending"] is False
def test_validation_is_unchanged_but_all_skips_it(self, web, socket_up):
assert web.post("/api/v3/errors/clear", json={"max_age_hours": 0}).status_code == 400
assert web.post("/api/v3/errors/clear", json={"max_age_hours": 9000}).status_code == 400
assert web.post("/api/v3/errors/clear",
json={"all": True, "max_age_hours": "junk"}).status_code == 200
class TestClearWithoutTheSocket:
"""Stage 5: no mailbox to fall back to. The route says why it failed and
writes nothing; the errors stay as the display last reported them."""
def _post(self, web, display, monkeypatch, error):
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(side_effect=error))
@pytest.mark.parametrize("reason", ["no_socket", "refused", "disabled", "unsupported"])
def test_no_socket_writes_the_mailbox(self, web, display, shared_cache, monkeypatch, reason):
from src.ipc import client as control_client
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(
side_effect=control_client.ControlError(reason, sent=False)))
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
return web.post("/api/v3/errors/clear", json={"all": True})
data = web.post("/api/v3/errors/clear", json={"all": True}).get_json()["data"]
assert data["transport"] == "mailbox" and data["applied"] is False
assert _mailbox_file(shared_cache).exists()
assert _summary(web)["clear_pending"] is True
publisher.tick()
assert aggregator.get_error_summary()["total_errors"] == 0
@pytest.mark.parametrize("reason", ["no_socket", "refused"])
def test_a_stopped_display_is_an_error(self, web, display, shared_cache, monkeypatch,
reason):
def test_an_older_display_gets_the_mailbox(self, web, display, shared_cache, monkeypatch):
# The upgrade case: new web interface, a display from before errors.clear.
from src.ipc import client as control_client
response = self._post(web, display, monkeypatch,
control_client.ControlError(reason, sent=False))
assert response.status_code == 503
body = response.get_json()
assert body["context"]["socket_error"] == reason
assert "not running" in body["message"]
assert not _mailbox_file(shared_cache).exists()
# Nothing hides them: they are still the display's last report.
summary = _summary(web)
assert summary["total_errors"] == 1 and summary["clear_pending"] is False
@pytest.mark.parametrize("reason", ["disabled", "unsupported"])
def test_no_socket_here_is_an_error(self, web, display, shared_cache, monkeypatch, reason):
from src.ipc import client as control_client
response = self._post(web, display, monkeypatch,
control_client.ControlError(reason, sent=False))
assert response.status_code == 503
assert "not available" in response.get_json()["message"]
assert not _mailbox_file(shared_cache).exists()
def test_an_older_display_is_told_to_restart(self, web, display, shared_cache,
monkeypatch):
from src.ipc import client as control_client
response = self._post(web, display, monkeypatch,
control_client.ControlError("unknown_command", sent=True))
assert response.status_code == 503
assert "restart" in response.get_json()["message"]
assert not _mailbox_file(shared_cache).exists()
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(
side_effect=control_client.ControlError("unknown_command", sent=True)))
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
data = web.post("/api/v3/errors/clear", json={"all": True}).get_json()["data"]
assert data["transport"] == "mailbox"
publisher.tick()
assert aggregator.get_error_summary()["total_errors"] == 0
@pytest.mark.parametrize("reason", ["internal", "timeout", "busy", "invalid_args"])
def test_a_display_that_had_it_and_failed_is_an_error(self, web, display, shared_cache,
monkeypatch, reason):
from src.ipc import client as control_client
response = self._post(web, display, monkeypatch,
control_client.ControlError(reason, sent=True))
assert response.status_code == 503
body = response.get_json()
assert body["context"]["socket_error"] == reason
assert body["message"] == "The display service did not apply the clear"
assert not _mailbox_file(shared_cache).exists()
def test_the_default_test_setup_has_no_socket(self, web, display, shared_cache):
# conftest turns the socket off: the real client answers "disabled".
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(
side_effect=control_client.ControlError(reason, sent=True)))
response = web.post("/api/v3/errors/clear", json={"all": True})
assert response.status_code == 503
assert response.get_json()["context"]["socket_error"] in ("disabled", "unsupported")
assert response.get_json()["context"]["socket_error"] == reason
assert not _mailbox_file(shared_cache).exists()
class TestPublisher:
def test_a_tick_reads_no_clear_request(self, display, shared_cache):
_, publisher, clock = display
_, web_cache, directory = shared_cache
# An old web interface's leftover request file is ignored.
(directory / f"{RETIRED_CLEAR_KEY}.json").write_text(
'{"timestamp": 1, "data": {"request_id": "old", "cutoff": 9e9}}')
class TestPublisherMailboxPoll:
def test_the_mailbox_is_read_only_when_its_file_changed(self, display, shared_cache):
_, publisher, _ = display
_, web_cache, _ = shared_cache
publisher.tick()
publisher.cache_manager = MagicMock(wraps=publisher.cache_manager)
def reads():
return [c for c in publisher.cache_manager.get.call_args_list
if c.args[0] == ERROR_CLEAR_REQUEST_KEY]
for _ in range(5):
clock.now += 60
publisher.tick()
publisher.cache_manager.get.assert_not_called()
snapshot = web_cache.get(ERROR_SNAPSHOT_KEY, max_age=None, memory_ttl=0)
assert snapshot["applied_clear_id"] is None
assert reads() == [] # no file: a stat per tick, no read
errors.request_error_clear(web_cache, 1.0)
publisher.tick()
assert len(reads()) == 1
for _ in range(5):
publisher.tick()
assert len(reads()) == 1 # unchanged file: not read again
errors.request_error_clear(web_cache, 2.0)
publisher.tick()
assert len(reads()) == 2
def test_clear_now_publishes_what_it_applied(self, display, shared_cache):
aggregator, publisher, _ = display
@@ -502,6 +572,7 @@ class TestPublisher:
snapshot = web_cache.get(ERROR_SNAPSHOT_KEY, max_age=None, memory_ttl=0)
assert snapshot["applied_clear_id"] == "sock-1"
assert snapshot["total_errors"] == 0
assert snapshot["applied_clear_cutoff"] is not None
def test_the_handler_needs_a_running_publisher(self, monkeypatch):
from src.ipc.contract import ErrorsClearArgs
@@ -512,10 +583,12 @@ class TestPublisher:
@pytest.mark.skipif(not hasattr(os, "fchmod") or os.name == "nt",
reason="POSIX file modes")
def test_the_snapshot_is_group_readable(web, display, shared_cache):
def test_both_files_are_group_readable(web, display, shared_cache):
aggregator, publisher, _ = display
_, _, directory = shared_cache
_fail(aggregator)
publisher.tick()
mode = stat.S_IMODE(os.stat(directory / f"{ERROR_SNAPSHOT_KEY}.json").st_mode)
assert mode == 0o660, oct(mode)
web.post("/api/v3/errors/clear", json={"all": True})
for key in (ERROR_SNAPSHOT_KEY, ERROR_CLEAR_REQUEST_KEY):
mode = stat.S_IMODE(os.stat(directory / f"{key}.json").st_mode)
assert mode == 0o660, (key, oct(mode))
+3 -3
View File
@@ -3,7 +3,7 @@
Pure data, so every test here runs on every platform. What they pin:
* a request and a response survive encode -> decode -> parse unchanged, and
the on-demand arguments carry exactly what the REST route sends;
the on-demand arguments carry exactly what the file mailbox carries;
* the envelope and the arguments refuse what the display could not act on
(missing ids, wrong types, a non-finite duration) with a stable error code;
* framing never holds more than one message's worth of bytes, however the
@@ -184,8 +184,8 @@ class TestOnDemandArgs:
class TestMailboxShape:
"""Socket commands are handed to the display's on-demand handler, so they
must look exactly like the request dict it takes."""
"""Socket commands are handed to the mailbox's own handler, so they must
look exactly like what the web route writes to the mailbox."""
def test_start(self):
args = OnDemandStartArgs(plugin_id='clock', mode='clock_main', duration=60.0,
+58 -34
View File
@@ -1,13 +1,14 @@
"""DisplayController's side of the control socket.
The server's handlers only queue; the render thread drains the queue
(_poll_on_demand_requests) and hands each on-demand command to
_handle_on_demand_request, which plugins' own requests use too. These tests
pin that hook:
The server's handlers only queue; the render thread drains the queue where
it reads the file mailbox (_poll_on_demand_requests) and hands each command
to the mailbox's own handler (_handle_on_demand_request). These tests pin
that hook:
* a socket command is applied with its request id, at once, and touches no
cache key (the file mailbox and the persisted processed id are gone);
* a start sent twice with one request id is activated once;
* a socket command is applied by the same code as a mailbox request, with
its request id, and without waiting for the mailbox's 0.25 s read floor;
* a request that arrives both ways (a client that timed out after the
command was queued, then wrote the mailbox) is activated once;
* a command that fails is contained, and the ones after it still run;
* cleanup closes the socket; a disabled socket changes nothing.
"""
@@ -57,15 +58,28 @@ def controller(test_display_controller):
c_ = test_display_controller
c_.on_demand_active = False
c_.on_demand_request_id = None
c_.cache_manager.get = MagicMock(return_value=None)
c_._last_on_demand_poll = None
mailbox = {'value': None}
def fake_get(key, *a, **kw):
if key == 'display_on_demand_request':
return mailbox['value']
return None
def fake_delete(key):
if key == 'display_on_demand_request':
mailbox['value'] = None
c_.cache_manager.get = MagicMock(side_effect=fake_get)
c_.cache_manager.set = MagicMock()
c_.cache_manager.delete = MagicMock()
c_.cache_manager.delete = MagicMock(side_effect=fake_delete)
c_._activate_on_demand = MagicMock()
c_.mailbox = mailbox
return c_
class TestDrain:
def test_a_socket_start_is_activated_with_its_request_id(self, controller):
def test_a_socket_start_goes_through_the_mailbox_handler(self, controller):
controller._control_server = FakeServer(_start('sock-1', duration=30.0, pinned=True))
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
@@ -75,28 +89,47 @@ class TestDrain:
assert request['plugin_id'] == 'clock'
assert request['duration'] == 30.0 and request['pinned'] is True
assert controller.on_demand_request_id == 'sock-1'
# No mailbox, and no persisted processed id.
keys = {call.args[0] for m in (controller.cache_manager.get,
controller.cache_manager.set,
controller.cache_manager.delete)
for call in m.call_args_list}
assert not keys & {'display_on_demand_request', 'display_on_demand_processed_id'}
# The same restart-replay guard as a mailbox request.
controller.cache_manager.set.assert_any_call(
'display_on_demand_processed_id', 'sock-1', ttl=3600)
def test_socket_commands_land_at_once(self, controller):
def test_socket_commands_skip_the_mailbox_floor(self, controller):
server = FakeServer()
controller._control_server = server
controller._poll_on_demand_requests()
controller._poll_on_demand_requests() # reads the mailbox, sets the floor
reads = controller.cache_manager.get.call_count
server.commands.append(_start('quick'))
controller._poll_on_demand_requests() # within the floor
controller._activate_on_demand.assert_called_once()
mailbox_reads = [call for call in controller.cache_manager.get.call_args_list[reads:]
if call.args[0] == 'display_on_demand_request']
# Only _consume_on_demand_request's compare-before-delete re-read.
assert len(mailbox_reads) <= 1
def test_a_request_that_came_both_ways_is_activated_once(self, controller):
controller._control_server = FakeServer(_start('dup'))
controller.mailbox['value'] = {'request_id': 'dup', 'action': 'start',
'plugin_id': 'clock'}
controller._poll_on_demand_requests()
controller._last_on_demand_poll = None
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
assert controller.mailbox['value'] is None, "the duplicate was left in the mailbox"
def test_a_fallback_write_landing_later_is_ignored(self, controller):
controller._control_server = FakeServer(_start('late'))
controller._poll_on_demand_requests()
controller.mailbox['value'] = {'request_id': 'late', 'action': 'start',
'plugin_id': 'clock'}
controller._last_on_demand_poll = None
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
def test_a_start_sent_twice_is_activated_once(self, controller):
server = FakeServer(_start('dup'))
controller._control_server = server
def test_the_mailbox_still_works_alongside(self, controller):
controller._control_server = FakeServer()
controller.mailbox['value'] = {'request_id': 'mb', 'action': 'start', 'plugin_id': 'p'}
controller._poll_on_demand_requests()
server.commands.append(_start('dup'))
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
assert controller._activate_on_demand.call_args.args[0]['request_id'] == 'mb'
def test_a_socket_stop_ends_on_demand(self, controller):
controller.on_demand_active = True
@@ -126,11 +159,10 @@ class TestDrain:
controller._poll_on_demand_requests()
assert calls == ['bad', 'good']
def test_no_server_means_nothing_to_apply(self, controller):
def test_no_server_means_mailbox_only(self, controller):
controller._control_server = None
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_not_called()
controller.cache_manager.get.assert_not_called()
class TestPendingChangesFloor:
@@ -162,14 +194,6 @@ class TestLifecycle:
assert status['on_demand']['plugin_id'] == 'clock'
c.encode_message(status) # it has to fit on the wire
def test_the_state_names_the_request_it_answers(self, controller):
# The web interface keeps reporting a delivered start as "starting"
# until the display publishes state for that request id.
assert controller._on_demand_state()['request_id'] is None
controller._control_server = FakeServer(_start('sock-7'))
controller._poll_on_demand_requests()
assert controller._on_demand_state()['request_id'] == 'sock-7'
def test_cleanup_closes_the_socket(self, controller):
server = FakeServer()
controller._control_server = server
+209 -144
View File
@@ -1,13 +1,14 @@
"""Stages 4 and 5 of the control socket: the socket is the only way in.
"""Stage 4 of the control socket: the file mailboxes are only a fallback.
* The client knows whether the display had the request (``ControlError.sent``)
and ``display_not_listening`` says when no display was there to take it
(stopped or still starting), the one case a later retry can fix.
and ``should_fall_back`` allows a mailbox write only when it did not, or
when the display is too old to know the command (the upgrade case).
* ``errors.clear`` is answered on the connection thread by a handler the
display registers; a display without one answers like an older display.
* Stage 5: the display no longer reads the ``display_on_demand_request``
mailbox at all, and the cache refuses writes to the retired mailbox keys,
warning once per writer.
* The display looks at the on-demand mailbox once a second while the socket
is up (0.25 s without it), reads it only when its file changed, never
touches it for a socket command, and logs who still writes it.
* ``CacheManager.file_signature`` / ``MailboxWatch`` make a look one stat().
The web routes are covered in test_api_v3_on_demand_socket.py and
test_error_snapshot_cross_process.py.
@@ -22,8 +23,7 @@ from unittest.mock import MagicMock, patch
import pytest
from src import cache_manager as cache_module
from src.cache_manager import CacheManager
from src.cache_manager import CacheManager, MailboxWatch
from src.ipc import client
from src.ipc import contract as c
from src.ipc.contract import Command, ErrorsClearArgs, OnDemandStartArgs, ProtocolError
@@ -81,41 +81,38 @@ class TestSent:
def test_an_answer_is_returned(self):
assert _call(FakeSock([_reply('rid-1', ok=True, result={'pong': True})])) == {'pong': True}
@pytest.mark.parametrize('reason,listening', [('no_socket', False), ('refused', False),
('timeout', True), ('busy', True)])
def test_a_failed_connect_was_not_sent(self, reason, listening):
@pytest.mark.parametrize('reason', ['no_socket', 'refused', 'timeout', 'busy'])
def test_a_failed_connect_was_not_sent(self, reason):
e = _error(connect_error=client.ControlError(reason))
assert e.reason == reason and e.sent is False
# Only "nothing there" is worth waiting for: a timeout or a full
# backlog is a display that is there and stuck.
assert client.display_not_listening(e) is not listening
assert client.should_fall_back(e)
def test_a_send_that_timed_out_was_not_sent(self):
e = _error(sock=FakeSock(send_error=socket.timeout()))
assert e.reason == 'timeout' and e.sent is False
assert not client.display_not_listening(e)
assert client.should_fall_back(e)
def test_silence_after_the_request_was_sent(self):
e = _error(sock=FakeSock(recv_error=socket.timeout()))
assert e.reason == 'timeout' and e.sent is True
assert not client.display_not_listening(e)
assert not client.should_fall_back(e)
def test_a_hang_up_after_the_request_was_sent(self):
e = _error(sock=FakeSock([]))
assert e.reason == 'closed' and e.sent is True
assert not client.display_not_listening(e)
assert not client.should_fall_back(e)
def test_a_garbled_reply(self):
e = _error(sock=FakeSock([b'not json\n']))
assert e.reason == 'bad_response' and e.sent is True
assert not client.display_not_listening(e)
assert not client.should_fall_back(e)
@pytest.mark.parametrize('code', ['busy', 'invalid_args', 'internal', 'pending', 'failed'])
def test_a_display_error_with_an_id_was_sent(self, code):
sock = FakeSock([_reply('rid-1', ok=False, error={'code': code, 'message': 'x'})])
e = _error(sock=sock)
assert e.reason == code and e.sent is True
assert not client.display_not_listening(e)
assert not client.should_fall_back(e)
@pytest.mark.parametrize('code', ['forbidden', 'busy'])
def test_a_refusal_at_the_door_was_not_sent(self, code):
@@ -125,33 +122,23 @@ class TestSent:
'error': {'code': code, 'message': 'x'}}) + '\n').encode()])
e = _error(sock=sock)
assert e.reason == code and e.sent is False
assert not client.display_not_listening(e)
assert client.should_fall_back(e)
@pytest.mark.parametrize('code', ['unknown_command', 'unsupported_version'])
def test_an_older_display_is_listening(self, code):
def test_an_older_display_is_fallen_back_from(self, code):
sock = FakeSock([_reply('rid-1', ok=False, error={'code': code, 'message': 'x'})])
e = _error(sock=sock)
assert e.sent is True
assert not client.display_not_listening(e)
assert client.should_fall_back(e)
def test_a_request_refused_locally_never_left(self):
with pytest.raises(client.ControlError) as e:
client.request(Command.ERRORS_CLEAR, {'cutoff': 'soon'}, paths=['/x.sock'])
assert e.value.reason == 'invalid_request' and e.value.sent is False
assert not client.display_not_listening(e.value)
assert client.should_fall_back(e.value)
@pytest.mark.parametrize('reason', ['disabled', 'unsupported'])
def test_a_client_without_the_socket_is_not_waiting_for_a_display(self, reason):
assert not client.display_not_listening(client.ControlError(reason))
@pytest.mark.parametrize('reason', sorted(client.NOT_LISTENING_REASONS))
def test_a_request_that_was_sent_had_a_display(self, reason):
# Only a request that never left can be waited on and sent again: one
# that was sent may have been applied.
assert not client.display_not_listening(client.ControlError(reason, sent=True))
def test_a_client_bug_is_not_a_missing_display(self):
assert not client.display_not_listening(RuntimeError('boom'))
def test_a_client_bug_falls_back(self):
assert client.should_fall_back(RuntimeError('boom'))
# -- errors.clear on the server ------------------------------------------------------
@@ -217,28 +204,37 @@ class TestErrorsClearOnTheServer:
assert Command.ERRORS_CLEAR in result['commands']
# -- stage 5: the display reads no mailbox --------------------------------------------
# -- the display's mailbox poll ------------------------------------------------------
class CountingCache:
"""The slice of CacheManager the on-demand path uses, counting every call."""
class SignedCache:
"""The slice of CacheManager the poll uses, counting what it costs."""
def __init__(self):
self.data = {}
self.calls = []
self.writes = 0
self.version = {}
self.reads = []
self.stats = 0
self.deletes = []
self.sets = []
def __getattr__(self, name):
# Any other method (file_signature, delete, ...) is recorded too.
def call(*a, **kw):
self.calls.append((name, a[0] if a else None))
return call
def file_signature(self, key):
self.stats += 1
return (self.version[key], 0, 0) if key in self.data else None
def get(self, key, *a, **kw):
self.calls.append(('get', key))
self.reads.append(key)
return self.data.get(key)
def set(self, key, value, *a, **kw):
self.calls.append(('set', key))
self.sets.append(key)
self.data[key] = value
self.writes += 1
self.version[key] = self.writes
def delete(self, key):
self.deletes.append(key)
self.data.pop(key, None)
class FakeServer:
@@ -254,140 +250,209 @@ class FakeServer:
return out
class Clock:
def __init__(self):
self.t = 1000.0
def __call__(self):
return self.t
@pytest.fixture
def controller(test_display_controller):
def controller(test_display_controller, monkeypatch):
dc = test_display_controller
dc.cache_manager = CountingCache()
dc.cache_manager = SignedCache()
dc._activate_on_demand = MagicMock()
dc.on_demand_active = False
dc.on_demand_request_id = None
dc._last_on_demand_poll = None
dc._on_demand_mailbox = None
dc._mailbox_writers_logged = frozenset()
clock = Clock()
monkeypatch.setattr('src.display_controller.time.monotonic', clock)
dc.clock = clock
return dc
class TestNoMailbox:
@pytest.mark.parametrize('with_socket', [False, True])
def test_polling_touches_no_cache_key(self, controller, with_socket):
controller._control_server = FakeServer() if with_socket else None
controller.cache_manager.data[MAILBOX] = {'request_id': 'left', 'action': 'start',
'plugin_id': 'clock'}
for _ in range(200):
controller._poll_on_demand_requests()
assert controller.cache_manager.calls == []
controller._activate_on_demand.assert_not_called()
def _post(dc, rid, action='start', **fields):
dc.cache_manager.set(MAILBOX, dict({'request_id': rid, 'action': action}, **fields))
def test_socket_commands_land_at_once(self, controller):
def _poll_for(dc, seconds, step=1 / 16): # exact in binary: no drift past a floor
end = dc.clock.t + seconds
while dc.clock.t < end:
dc._poll_on_demand_requests()
dc.clock.t += step
class TestMailboxCadence:
def test_without_a_socket_it_is_looked_at_every_quarter_second(self, controller):
controller._control_server = None
_poll_for(controller, 10.0)
assert 38 <= controller.cache_manager.stats <= 42
def test_with_a_socket_it_is_looked_at_once_a_second(self, controller):
controller._control_server = FakeServer()
_poll_for(controller, 10.0)
assert 9 <= controller.cache_manager.stats <= 11
def test_a_look_that_finds_nothing_reads_nothing(self, controller):
controller._control_server = FakeServer()
_poll_for(controller, 10.0)
assert controller.cache_manager.reads == []
def test_an_unchanged_mailbox_is_not_read_again(self, controller):
# An already-processed start the delete could not remove, say.
controller._control_server = FakeServer()
controller.cache_manager.delete = MagicMock() # the file stays
_post(controller, 'once', plugin_id='clock')
_poll_for(controller, 10.0)
assert controller.cache_manager.reads.count(MAILBOX) <= 2 # the read + the re-check
controller._activate_on_demand.assert_called_once()
def test_a_mailbox_request_lands_within_a_second_with_the_socket_up(self, controller):
# The upgrade case the other way round: a new display, and a web
# interface (or a plugin) that still writes the mailbox.
controller._control_server = FakeServer()
controller._poll_on_demand_requests()
controller.clock.t += 0.1
_post(controller, 'old-web', plugin_id='clock')
posted = controller.clock.t
while not controller._activate_on_demand.called:
controller._poll_on_demand_requests()
controller.clock.t += 0.05
assert controller.clock.t - posted < 1.5
assert controller.clock.t - posted <= controller.MAILBOX_POLL_INTERVAL_WITH_SOCKET + 0.06
assert MAILBOX in controller.cache_manager.deletes # consumed
def test_socket_commands_still_land_at_once(self, controller):
server = controller._control_server = FakeServer()
controller._poll_on_demand_requests()
server.commands.append(QueuedCommand('sock', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests() # inside the mailbox interval
controller._activate_on_demand.assert_called_once()
class TestSocketCommandsLeaveTheMailboxAlone:
def test_a_socket_start_reads_and_deletes_no_mailbox(self, controller):
server = controller._control_server = FakeServer()
controller._poll_on_demand_requests()
before = list(controller.cache_manager.reads)
server.commands.append(QueuedCommand('s1', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
assert MAILBOX not in controller.cache_manager.reads[len(before):]
assert controller.cache_manager.deletes == []
def test_a_start_is_applied_once_and_writes_no_processed_id(self, controller):
server = controller._control_server = FakeServer()
for _ in range(2):
server.commands.append(QueuedCommand('same', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
assert controller.cache_manager.calls == []
def test_a_socket_stop_touches_no_cache_key(self, controller):
def test_a_socket_stop_reads_and_deletes_no_mailbox(self, controller):
from src.ipc.contract import OnDemandStopArgs
controller.on_demand_active = True
controller._clear_on_demand = MagicMock()
server = controller._control_server = FakeServer()
server.commands.append(QueuedCommand('s2', Command.ON_DEMAND_STOP,
OnDemandStopArgs(), time.time()))
controller.clock.t += 5
controller._drain_control_commands()
controller._clear_on_demand.assert_called_once()
assert controller.cache_manager.calls == []
assert MAILBOX not in controller.cache_manager.reads
assert controller.cache_manager.deletes == []
def test_the_mailbox_helpers_are_gone(self):
from src import display_controller as dcm
from src import error_aggregator as ea
for name in ('ON_DEMAND_MAILBOX_KEY', 'MailboxWatch'):
assert not hasattr(dcm, name)
assert not hasattr(ea, 'ERROR_CLEAR_REQUEST_KEY')
assert not hasattr(cache_module, 'MailboxWatch')
assert not hasattr(CacheManager, 'file_signature')
assert not hasattr(client, 'should_fall_back')
for name in ('_consume_on_demand_request', '_note_mailbox_request',
'_mailbox_poll_interval', 'MAILBOX_POLL_INTERVAL_WITH_SOCKET'):
assert not hasattr(dcm.DisplayController, name)
def test_a_mailbox_copy_of_a_socket_command_is_dropped(self, controller):
# An older web interface timed out after the display queued the
# command, then wrote the mailbox too.
server = controller._control_server = FakeServer()
server.commands.append(QueuedCommand('both', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests()
_post(controller, 'both', plugin_id='clock')
_poll_for(controller, 2.0)
controller._activate_on_demand.assert_called_once()
assert MAILBOX not in controller.cache_manager.data
# -- stage 5: writes to the retired keys are refused, with one warning per writer -----
class TestDeprecationLog:
def test_each_mailbox_writer_is_logged_once(self, controller, caplog):
controller._control_server = FakeServer()
caplog.set_level(logging.INFO, logger='src.display_controller')
for i, plugin in enumerate(['on-air', 'on-air', 'pomodoro-timer']):
_post(controller, f'r{i}', plugin_id=plugin)
_poll_for(controller, 1.2)
lines = [r.getMessage() for r in caplog.records if 'file mailbox' in r.getMessage()]
assert len(lines) == 2
assert 'on-air' in lines[0] and 'pomodoro-timer' in lines[1]
def test_nothing_is_logged_without_a_socket(self, controller, caplog):
controller._control_server = None
caplog.set_level(logging.INFO, logger='src.display_controller')
_post(controller, 'r', plugin_id='on-air')
_poll_for(controller, 1.0)
controller._activate_on_demand.assert_called_once()
assert not [r for r in caplog.records if 'file mailbox' in r.getMessage()]
# -- file_signature and MailboxWatch -------------------------------------------------
@pytest.fixture
def real_cache(tmp_path, monkeypatch):
monkeypatch.setattr(CacheManager, '_get_writable_cache_dir', lambda self: str(tmp_path))
monkeypatch.setattr(cache_module, '_retired_writers_warned', set())
cache = CacheManager()
yield cache
cache.stop_cleanup_thread()
class OldPlugin:
"""What an old plugin looks like on the stack: BasePlugin gives every
plugin ``plugin_id`` and ``cache_manager``."""
class TestFileSignature:
def test_absent_key(self, real_cache):
assert real_cache.file_signature('nothing') is None
def __init__(self, plugin_id, cache):
self.plugin_id = plugin_id
self.cache_manager = cache
def test_every_write_is_a_new_signature(self, real_cache):
seen = set()
for i in range(20):
# Same size each time, written as fast as possible.
real_cache.set(MAILBOX, {'request_id': f'r{i:02d}'})
sig = real_cache.file_signature(MAILBOX)
assert isinstance(sig, tuple)
seen.add(sig)
assert len(seen) == 20
def trigger(self, target=None):
self.cache_manager.set(MAILBOX, {'request_id': 'r', 'action': 'start',
'plugin_id': target or self.plugin_id})
def test_gone_after_a_delete(self, real_cache):
real_cache.set(MAILBOX, {'a': 1})
real_cache.delete(MAILBOX)
assert real_cache.file_signature(MAILBOX) is None
def _warnings(caplog):
return [r.getMessage() for r in caplog.records
if r.levelno == logging.WARNING and 'retired' in r.getMessage()]
class TestMailboxWatch:
def test_reads_once_per_write(self, real_cache):
watch = MailboxWatch(MAILBOX)
assert watch.changed(real_cache) is False # no file
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache) is True
assert watch.changed(real_cache) is False
real_cache.set(MAILBOX, {'request_id': 'b'})
assert watch.changed(real_cache) is True
def test_forget_reads_again(self, real_cache):
watch = MailboxWatch(MAILBOX)
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache) is True
watch.forget()
assert watch.changed(real_cache) is True
class TestRetiredKeys:
@pytest.mark.parametrize('key', sorted(cache_module.RETIRED_MAILBOX_KEYS))
def test_a_write_stores_nothing(self, real_cache, key):
real_cache.set(key, {'request_id': 'r'})
real_cache.save_cache(key, {'request_id': 'r2'})
assert real_cache.get(key, max_age=None, memory_ttl=0) is None
assert not [f for f in os.listdir(real_cache.cache_dir) if key in f]
def test_a_rewrite_after_a_delete_is_seen(self, real_cache):
watch = MailboxWatch(MAILBOX)
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache)
real_cache.delete(MAILBOX)
assert watch.changed(real_cache) is False
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache) is True
def test_other_keys_are_unaffected(self, real_cache):
real_cache.set('display_on_demand_state', {'active': True})
assert real_cache.get('display_on_demand_state', max_age=None,
memory_ttl=0) == {'active': True}
def test_the_writing_plugin_is_named_once(self, real_cache, caplog):
caplog.set_level(logging.WARNING)
on_air = OldPlugin('on-air', real_cache)
for _ in range(3):
on_air.trigger()
OldPlugin('pomodoro-timer', real_cache).trigger()
lines = _warnings(caplog)
assert len(lines) == 2
assert "plugin 'on-air'" in lines[0] and 'display_on_demand_request' in lines[0]
assert 'request_on_demand' in lines[0]
assert "plugin 'pomodoro-timer'" in lines[1]
def test_the_writer_is_the_caller_not_the_target(self, real_cache, caplog):
caplog.set_level(logging.WARNING)
OldPlugin('mqtt-notifications', real_cache).trigger(target='clock')
(line,) = _warnings(caplog)
assert "plugin 'mqtt-notifications'" in line and 'clock' not in line
def test_without_a_plugin_on_the_stack_the_request_names_it(self, real_cache, caplog):
caplog.set_level(logging.WARNING)
real_cache.set(MAILBOX, {'request_id': 'r', 'action': 'start', 'plugin_id': 'gif-player'})
(line,) = _warnings(caplog)
assert "plugin 'gif-player' (named in the request)" in line
def test_otherwise_unknown(self, real_cache, caplog):
caplog.set_level(logging.WARNING)
real_cache.set('plugin_error_clear_request', {'request_id': 'r', 'cutoff': 1.0})
real_cache.set('plugin_error_clear_request', {'request_id': 'r2', 'cutoff': 2.0})
(line,) = _warnings(caplog)
assert 'by unknown' in line and '/api/v3/errors/clear' in line
def test_a_cache_that_cannot_tell_is_read_every_time(self):
watch = MailboxWatch(MAILBOX)
assert watch.changed(MagicMock()) is True
assert watch.changed(MagicMock()) is True
assert watch.changed(object()) is True
# -- end to end over a real socket ---------------------------------------------------
@@ -414,24 +479,24 @@ class TestOverTheSocket:
finally:
server.close()
def test_an_older_display_is_listening(self, sock_path):
def test_an_older_display_is_an_upgrade_fallback(self, sock_path):
server = ControlServer(sock_path) # no errors.clear handler
assert server.start()
try:
with pytest.raises(client.ControlError) as e:
client.errors_clear('clr-2', 1.0, paths=[sock_path])
assert e.value.reason == 'unknown_command' and e.value.sent is True
assert not client.display_not_listening(e.value)
assert client.should_fall_back(e.value)
finally:
server.close()
def test_no_display_is_not_listening(self, sock_path):
def test_no_display_is_a_fallback(self, sock_path):
with pytest.raises(client.ControlError) as e:
client.errors_clear('clr-3', 1.0, paths=[sock_path])
assert e.value.reason == 'no_socket' and e.value.sent is False
assert client.display_not_listening(e.value)
assert client.should_fall_back(e.value)
def test_a_full_queue_is_listening(self, sock_path):
def test_a_full_queue_is_not_a_fallback(self, sock_path):
server = ControlServer(sock_path, queue_size=1)
assert server.start()
try:
@@ -439,6 +504,6 @@ class TestOverTheSocket:
with pytest.raises(client.ControlError) as e:
client.on_demand_start('q2', 'clock', None, paths=[sock_path])
assert e.value.reason == 'busy' and e.value.sent is True
assert not client.display_not_listening(e.value)
assert not client.should_fall_back(e.value)
finally:
server.close()
-24
View File
@@ -327,27 +327,3 @@ class TestCleartextIsCalledOut:
def test_tls_on_does_not_warn(self, bridge_module):
assert bridge_module.warn_if_cleartext(
{"mqtt_tls": True, "mqtt_password": "hunter2"}) is False
class TestTheApiClientReadsStartingAsTaken:
"""A cold start answers 202 with ``status: "starting"``: the request is
taken and the web process delivers it once the display listens. The
bridge must report that as success, not as a failure."""
def _client(self, bridge_module, status_code, body):
response = MagicMock(status_code=status_code)
response.json.return_value = body
session = MagicMock()
session.request.return_value = response
return bridge_module.LEDMatrixClient("http://pi:5000", session=session)
def test_202_starting_is_returned_not_raised(self, bridge_module):
api = self._client(bridge_module, 202, {
"status": "starting", "message": "starting",
"data": {"request_id": "r1", "pending": True}})
assert api.start_on_demand(mode="clock") == {"request_id": "r1", "pending": True}
def test_an_error_is_still_raised(self, bridge_module):
api = self._client(bridge_module, 503, {"status": "error", "message": "no display"})
with pytest.raises(RuntimeError, match="no display"):
api.start_on_demand(mode="clock")
+7 -2
View File
@@ -337,8 +337,13 @@ class TestResumingAfterTheSession:
class TestStopClearsAnError:
def _post_stop(self, c):
c._handle_on_demand_request({'request_id': 'S1', 'action': 'stop',
'source': 'socket'})
stop = {'request_id': 'S1', 'action': 'stop'}
c._last_on_demand_poll = None
c.cache_manager.get = MagicMock(
side_effect=lambda key, *a, **kw:
stop if key == 'display_on_demand_request' else None)
c.cache_manager.delete = MagicMock()
c._poll_on_demand_requests()
def test_a_stop_after_a_failed_request_clears_the_error(self, controller):
_start(controller, plugin_id='uninstalled')
-272
View File
@@ -1,272 +0,0 @@
"""The web process's on-demand dispatcher (web_interface/on_demand_dispatch.py).
A start that finds no display listening -- the service was just started, or
is still loading its plugins -- is answered at once (202), and the
dispatcher's one worker thread sends it again until the display
acknowledges it or the wait runs out. These tests drive the worker with a
fake ``send`` and short waits:
* acknowledged: delivered once, and reported as such;
* nothing listening for the whole wait: ``start-timeout``;
* any other failure: reported at once, not retried;
* a newer start supersedes the pending one; a stop cancels it;
* the outcome is reported for a while, then forgotten.
"""
import threading
import time
import pytest
from src.ipc import client as control_client
from web_interface import on_demand_dispatch
from web_interface.on_demand_dispatch import OnDemandDispatcher
def _not_listening():
return control_client.ControlError("no_socket", "x", sent=False)
def _payload(rid, plugin_id="weather"):
return {"request_id": rid, "action": "start", "plugin_id": plugin_id,
"mode": plugin_id, "duration": 30, "pinned": False}
class FakeSend:
"""Answers with ``outcomes`` in turn (an exception is raised, anything
else acks); the last one repeats. Records the request ids it was sent."""
def __init__(self, *outcomes):
self.outcomes = list(outcomes) or ["ack"]
self.sent = []
self.lock = threading.Lock()
def __call__(self, payload):
with self.lock:
self.sent.append(payload["request_id"])
outcome = self.outcomes.pop(0) if len(self.outcomes) > 1 else self.outcomes[0]
if isinstance(outcome, BaseException):
raise outcome
if callable(outcome):
return outcome(payload)
return {"accepted": True}
def _until(predicate, timeout=5.0):
end = time.monotonic() + timeout
while time.monotonic() < end:
if predicate():
return True
time.sleep(0.005)
return False
def _settled(d):
return _until(lambda: not d.pending() and d._thread is None)
@pytest.fixture
def make():
made = []
def build(send, wait_seconds=2.0, retry_interval=0.01):
d = OnDemandDispatcher(send, wait_seconds=wait_seconds, retry_interval=retry_interval)
made.append(d)
return d
yield build
for d in made:
d.cancel("teardown")
class TestDelivery:
def test_an_acknowledged_start_is_delivered_once(self, make):
send = FakeSend("ack")
d = make(send)
d.submit(_payload("r1"))
assert _settled(d)
assert send.sent == ["r1"]
status = d.status()
assert status["status"] == "delivered" and status["request_id"] == "r1"
assert status["source"] == "web"
def test_it_is_sent_again_until_the_display_listens(self, make):
send = FakeSend(_not_listening(), _not_listening(), "ack")
d = make(send)
d.submit(_payload("r1"))
assert _settled(d)
assert send.sent == ["r1", "r1", "r1"]
assert d.status()["status"] == "delivered"
def test_while_it_waits_it_reports_starting(self, make):
gate = threading.Event()
def blocked(payload):
gate.wait(5)
raise _not_listening()
send = FakeSend(blocked, "ack")
d = make(send)
d.submit(_payload("r1"))
status = d.status()
assert status["status"] == "starting" and status["active"] is False
assert (status["plugin_id"], status["mode"], status["duration"]) == ("weather", "weather", 30)
gate.set()
assert _settled(d)
assert d.status()["status"] == "delivered"
class TestGivingUp:
def test_nothing_listening_for_the_whole_wait_is_a_start_timeout(self, make):
send = FakeSend(_not_listening())
d = make(send, wait_seconds=0.2)
started = time.monotonic()
d.submit(_payload("r1"))
assert _settled(d)
waited = time.monotonic() - started
status = d.status()
assert status["status"] == "error" and status["error"] == "start-timeout"
assert 0.15 <= waited < 2.0
assert len(send.sent) > 3 # it kept trying in between
def test_a_start_can_carry_its_own_wait(self, make):
# The route passes a shorter wait for a service that was already
# running; the dispatcher's default must not override it.
d = make(FakeSend(_not_listening()), wait_seconds=30.0)
d.submit(_payload("r1"), wait_seconds=0.1)
assert _until(lambda: not d.pending(), timeout=3.0), "it waited the default"
assert d.status()["error"] == "start-timeout"
@pytest.mark.parametrize("reason,sent", [("busy", True), ("unknown_command", True),
("timeout", True), ("forbidden", False)])
def test_any_other_failure_is_reported_at_once(self, make, reason, sent):
send = FakeSend(control_client.ControlError(reason, "x", sent=sent))
d = make(send)
d.submit(_payload("r1"))
assert _settled(d)
assert send.sent == ["r1"]
status = d.status()
assert status["status"] == "error" and status["error"] == reason
def test_a_client_bug_is_internal(self, make):
d = make(FakeSend(RuntimeError("boom")))
d.submit(_payload("r1"))
assert _settled(d)
assert d.status()["error"] == "internal"
class TestOneAtATime:
def test_a_newer_start_supersedes_the_pending_one(self, make):
send = FakeSend(_not_listening())
d = make(send)
d.submit(_payload("old"))
assert _until(lambda: "old" in send.sent)
d.submit(_payload("new", plugin_id="clock"))
assert d.status()["request_id"] == "new"
n = len(send.sent)
send.outcomes = ["ack"]
assert _settled(d)
assert send.sent[n:] and set(send.sent[n + 1:]) <= {"new"}
assert send.sent[-1] == "new"
status = d.status()
assert status["status"] == "delivered" and status["plugin_id"] == "clock"
def test_an_answer_for_a_superseded_start_is_not_reported(self, make):
# The old start's send is in flight when the new one arrives; its
# ack must not mark the new one delivered.
in_flight, release = threading.Event(), threading.Event()
def slow(payload):
in_flight.set()
release.wait(5)
return {"accepted": True}
send = FakeSend(slow, _not_listening())
d = make(send, wait_seconds=0.3)
d.submit(_payload("old"))
assert in_flight.wait(5)
d.submit(_payload("new"))
release.set()
assert _settled(d)
status = d.status()
assert status["request_id"] == "new"
assert status["status"] == "error" and status["error"] == "start-timeout"
def test_a_stop_cancels_the_pending_start(self, make):
send = FakeSend(_not_listening())
d = make(send)
d.submit(_payload("r1"))
assert _until(lambda: send.sent)
assert d.cancel("requested-stop") == "r1"
assert _settled(d)
n = len(send.sent)
time.sleep(0.05)
assert len(send.sent) == n, "it kept sending a cancelled start"
status = d.status()
assert status["status"] == "idle" and status["last_event"] == "requested-stop"
def test_a_cancel_waits_for_a_send_in_flight(self, make):
# Whatever the caller sends after cancel() must land after the
# cancelled start, not race it to the display.
in_flight, release = threading.Event(), threading.Event()
def slow(payload):
in_flight.set()
release.wait(5)
raise _not_listening()
send = FakeSend(slow)
d = make(send)
d.submit(_payload("old"))
assert in_flight.wait(5)
done = threading.Event()
threading.Thread(target=lambda: (d.cancel("superseded"), done.set()),
daemon=True).start()
assert not done.wait(0.1), "cancel returned while the old send was in flight"
release.set()
assert done.wait(5)
assert _settled(d)
assert send.sent == ["old"]
def test_a_cancel_with_nothing_pending_does_nothing(self, make):
d = make(FakeSend("ack"))
assert d.cancel() is None
assert d.status() is None
def test_a_cancel_after_delivery_leaves_the_outcome(self, make):
d = make(FakeSend("ack"))
d.submit(_payload("r1"))
assert _settled(d)
assert d.cancel() is None
assert d.status()["status"] == "delivered"
class TestOutcomeLifetime:
def test_an_outcome_is_forgotten_after_a_while(self, make, monkeypatch):
d = make(FakeSend("ack"))
d.submit(_payload("r1"))
assert _settled(d)
assert d.status() is not None
monkeypatch.setattr(on_demand_dispatch, "OUTCOME_SECONDS", 0.0)
time.sleep(0.01)
assert d.status() is None
def test_a_pending_start_never_expires(self, make, monkeypatch):
monkeypatch.setattr(on_demand_dispatch, "OUTCOME_SECONDS", 0.0)
gate = threading.Event()
d = make(FakeSend(lambda p: gate.wait(5) and {"accepted": True}))
d.submit(_payload("r1"))
time.sleep(0.01)
assert d.status()["status"] == "starting"
gate.set()
assert _settled(d)
def test_the_process_has_one_dispatcher():
on_demand_dispatch.reset_for_tests()
try:
assert on_demand_dispatch.current() is None
first = on_demand_dispatch.get_dispatcher(FakeSend())
assert on_demand_dispatch.get_dispatcher(FakeSend()) is first
assert on_demand_dispatch.current() is first
finally:
on_demand_dispatch.reset_for_tests()
+106
View File
@@ -0,0 +1,106 @@
"""The on-demand request mailbox: how often it is read, and how it is consumed.
The mailbox is a cache key the web process writes and the display process
reads. Two properties matter and neither is obvious from the call site:
* it is polled after every rendered frame, so an uncached read here is a
disk read at frame rate;
* consuming it must not throw away a request that arrived while the previous
one was being processed.
"""
from unittest.mock import MagicMock
import pytest
class TestPollingIsBounded:
"""_poll_on_demand_requests runs ~125x/second on a scrolling mode.
The read is deliberately uncached (memory_ttl=0) because a cached one
pinned the first request for an hour. That makes the call a real disk read,
so it needs a floor -- without one it was ~125 reads per second to find
nothing at all.
"""
def test_first_call_always_reads(self, test_display_controller):
c = test_display_controller
c.cache_manager.get = MagicMock(return_value=None)
c._poll_on_demand_requests()
assert c.cache_manager.get.call_count == 1
def test_immediate_second_call_does_not_read(self, test_display_controller):
c = test_display_controller
c.cache_manager.get = MagicMock(return_value=None)
c._poll_on_demand_requests()
for _ in range(50):
c._poll_on_demand_requests()
assert c.cache_manager.get.call_count == 1, "polling was not bounded"
def test_reads_again_once_the_interval_has_passed(self, test_display_controller, monkeypatch):
c = test_display_controller
c.cache_manager.get = MagicMock(return_value=None)
clock = {"t": 1000.0}
monkeypatch.setattr("src.display_controller.time.monotonic", lambda: clock["t"])
c._poll_on_demand_requests()
clock["t"] += c.ON_DEMAND_POLL_INTERVAL / 2
c._poll_on_demand_requests()
assert c.cache_manager.get.call_count == 1, "read before the interval elapsed"
clock["t"] += c.ON_DEMAND_POLL_INTERVAL
c._poll_on_demand_requests()
assert c.cache_manager.get.call_count == 2
def test_the_interval_is_short_enough_to_feel_instant(self, test_display_controller):
# A person clicking in the web UI must not notice the floor.
assert test_display_controller.ON_DEMAND_POLL_INTERVAL <= 0.5
class TestMailboxIsConsumedByIdentity:
"""Deleting whatever is in the mailbox loses a request that raced in."""
def _arrange(self, controller, first, later):
"""Mailbox returns `first`, then `later` on the pre-delete re-read."""
controller.on_demand_active = False
controller.on_demand_request_id = None
controller._last_on_demand_poll = None
reads = iter([first, later])
def fake_get(key, *a, **kw):
if key == 'display_on_demand_request':
return next(reads, later)
return None # processed-id lookup
controller.cache_manager.get = MagicMock(side_effect=fake_get)
controller.cache_manager.set = MagicMock()
controller.cache_manager.delete = MagicMock()
controller._activate_on_demand = MagicMock()
REQ_A = {'request_id': 'A', 'action': 'start', 'plugin_id': 'p', 'mode': 'm'}
REQ_B = {'request_id': 'B', 'action': 'start', 'plugin_id': 'p', 'mode': 'm'}
def test_own_request_is_deleted(self, test_display_controller):
c = test_display_controller
self._arrange(c, self.REQ_A, self.REQ_A)
c._poll_on_demand_requests()
c.cache_manager.delete.assert_called_once_with('display_on_demand_request')
def test_a_newer_request_is_left_for_the_next_poll(self, test_display_controller):
c = test_display_controller
self._arrange(c, self.REQ_A, self.REQ_B)
c._poll_on_demand_requests()
assert c.cache_manager.delete.call_count == 0, \
"request B was deleted without ever being processed"
def test_an_already_empty_mailbox_is_still_cleared(self, test_display_controller):
c = test_display_controller
self._arrange(c, self.REQ_A, None)
c._poll_on_demand_requests()
c.cache_manager.delete.assert_called_once_with('display_on_demand_request')
def test_the_request_is_still_processed(self, test_display_controller):
c = test_display_controller
self._arrange(c, self.REQ_A, self.REQ_B)
c._poll_on_demand_requests()
c._activate_on_demand.assert_called_once()
+38 -13
View File
@@ -8,9 +8,8 @@ Three separate gaps, all reachable from the web UI's force-display dialog:
* restarting while on-demand was active loaded *only* the on-demand plugin,
so normal rotation had nothing to return to for the life of the process;
* a stop request was exempt from the duplicate guards on purpose and was
never removed from the file mailbox, so it was re-processed on every poll
forever. The mailbox is gone (stage 5); a stop still skips the guards,
so a second click stops a session a race left running.
never removed from the mailbox, so it was re-processed on every poll
forever.
"""
from unittest.mock import MagicMock
@@ -186,25 +185,51 @@ class TestRestartDoesNotStarveTheOtherPlugins:
assert controller.on_demand_active is False
class TestStopRequestsSkipTheDuplicateGuard:
"""A stop is exempt from the request-id guard: every one is acted on."""
class TestStopRequestsAreConsumed:
"""A stop request is exempt from the duplicate guards, so the mailbox
delete is the only thing that ends it."""
STOP = {'request_id': 'S1', 'action': 'stop'}
def _arrange(self, controller, active):
controller.on_demand_active = active
controller.on_demand_status = 'active' if active else 'idle'
controller._last_on_demand_poll = None
controller.cache_manager.get = MagicMock(
side_effect=lambda key, *a, **kw:
self.STOP if key == 'display_on_demand_request' else None)
controller.cache_manager.set = MagicMock()
controller.cache_manager.delete = MagicMock()
controller._clear_on_demand = MagicMock()
def test_the_stop_is_acted_on(self, test_display_controller):
def test_a_handled_stop_is_removed_from_the_mailbox(self, test_display_controller):
c = test_display_controller
self._arrange(c, active=True)
c._handle_on_demand_request({'request_id': 'S1', 'action': 'stop',
'source': 'socket'})
c._poll_on_demand_requests()
c.cache_manager.delete.assert_called_once_with('display_on_demand_request')
def test_a_stop_arriving_while_idle_is_also_removed(self, test_display_controller):
"""Otherwise a stop sent to an idle display re-fires forever."""
c = test_display_controller
self._arrange(c, active=False)
c._poll_on_demand_requests()
c.cache_manager.delete.assert_called_once_with('display_on_demand_request')
def test_the_stop_is_still_acted_on(self, test_display_controller):
c = test_display_controller
self._arrange(c, active=True)
c._poll_on_demand_requests()
c._clear_on_demand.assert_called_once_with(reason='requested-stop')
def test_the_same_stop_twice_is_acted_on_twice(self, test_display_controller):
def test_a_start_racing_in_behind_a_stop_is_not_discarded(self, test_display_controller):
"""The compare-before-delete applies to stops too."""
c = test_display_controller
self._arrange(c, active=True)
for _ in range(2):
c._handle_on_demand_request({'request_id': 'S1', 'action': 'stop',
'source': 'socket'})
assert c._clear_on_demand.call_count == 2
newer = {'request_id': 'S2', 'action': 'start', 'plugin_id': 'p', 'mode': 'm'}
reads = iter([self.STOP, newer])
c.cache_manager.get = MagicMock(
side_effect=lambda key, *a, **kw:
next(reads, newer) if key == 'display_on_demand_request' else None)
c._poll_on_demand_requests()
assert c.cache_manager.delete.call_count == 0
+50 -38
View File
@@ -2,17 +2,19 @@
and end_on_demand().
A plugin running in the display process used to write the
``display_on_demand_request`` mailbox, which the display no longer reads
(stage 5). These tests pin the way in that replaced it:
``display_on_demand_request`` mailbox, which the display reads once a second
while the control socket is up. These tests pin the way in that replaces it:
* BasePlugin -> PluginManager -> DisplayController.submit_plugin_on_demand,
which only queues, from any thread;
* the render thread applies the queue where it applies socket commands,
through _handle_on_demand_request, at once, and woken by the control
socket when it is up;
through the mailbox's own handler, without the mailbox's read floor, and
woken by the control socket when it is up;
* a plugin's stop ends only its own session;
* no display to ask (the web interface's plugin manager, an old core's
plugin manager) answers None.
plugin manager) answers None, which is a plugin's cue to fall back to the
mailbox;
* the mailbox still works for plugins that write it.
"""
import logging
@@ -71,10 +73,19 @@ def controller(test_display_controller):
c_ = test_display_controller
c_.on_demand_active = False
c_.on_demand_request_id = None
c_.cache_manager.get = MagicMock(return_value=None)
c_._last_on_demand_poll = None
mailbox = {'value': None}
def fake_get(key, *a, **kw):
if key == 'display_on_demand_request':
return mailbox['value']
return None
c_.cache_manager.get = MagicMock(side_effect=fake_get)
c_.cache_manager.set = MagicMock()
c_.cache_manager.delete = MagicMock()
c_._activate_on_demand = MagicMock()
c_.mailbox = mailbox
return c_
@@ -90,7 +101,7 @@ class TestWiring:
controller.plugin_manager.set_on_demand_handler.assert_called_once_with(
controller.submit_plugin_on_demand)
def test_a_start_reaches_the_on_demand_handler(self, wired):
def test_a_start_reaches_the_mailbox_handler(self, wired):
controller, manager = wired
rid = _plugin('pomodoro-timer', manager).request_on_demand(
mode='pomodoro', duration=30, pinned=True)
@@ -128,6 +139,15 @@ class TestWiring:
controller._poll_on_demand_requests()
assert seen == ['a', 'b', 'c']
def test_the_mailbox_still_works_for_older_plugins(self, wired):
controller, manager = wired
controller.mailbox['value'] = {'request_id': 'mb', 'action': 'start',
'plugin_id': 'birdnet-go'}
_plugin('on-air', manager).request_on_demand()
controller._poll_on_demand_requests()
ids = [call.args[0]['request_id'] for call in controller._activate_on_demand.call_args_list]
assert 'mb' in ids and len(ids) == 2
def test_a_failing_request_is_contained(self, wired):
controller, manager = wired
calls = []
@@ -154,10 +174,10 @@ class TestPromptness:
controller._service_pending_changes() # well inside the 0.25 s floor
controller._activate_on_demand.assert_called_once()
def test_a_plugin_request_lands_on_the_next_poll(self, wired):
def test_a_plugin_request_skips_the_mailbox_floor(self, wired):
controller, manager = wired
controller._control_server = _WakeServer()
controller._poll_on_demand_requests()
controller._poll_on_demand_requests() # sets the 1 s mailbox floor
_plugin('p', manager).request_on_demand()
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
@@ -266,13 +286,14 @@ class TestStop:
controller._poll_on_demand_requests()
controller._clear_on_demand.assert_not_called()
def test_a_socket_stop_still_ends_any_session(self, wired):
def test_a_mailbox_stop_still_ends_any_session(self, wired):
controller, _ = wired
controller.on_demand_active = True
controller.on_demand_plugin_id = 'clock'
controller._clear_on_demand = MagicMock()
controller._handle_on_demand_request({'request_id': 's', 'action': 'stop',
'source': 'socket'})
controller.mailbox['value'] = {'request_id': 's', 'action': 'stop',
'plugin_id': 'on-air'}
controller._poll_on_demand_requests()
controller._clear_on_demand.assert_called_once_with(reason='requested-stop')
def test_start_then_stop_from_one_thread_ends_the_session(self, wired):
@@ -293,7 +314,7 @@ class TestStop:
class TestNoDisplay:
"""None: no display in this process took the request."""
"""None is a plugin's cue to write the mailbox instead."""
def test_a_manager_with_no_handler_answers_none(self):
plugin = _plugin('p', _manager())
@@ -332,30 +353,21 @@ class TestNoDisplay:
assert bare._plugin_on_demand_pending() is False
bare._drain_plugin_on_demand() # nothing to do, no error
def test_a_mailbox_write_after_none_is_dropped_with_a_warning(self, tmp_path,
monkeypatch, caplog):
"""What a plugin written for older cores does on None now: its
fallback write to the retired key stores nothing, and the log names
it once."""
from src import cache_manager as cache_module
from src.cache_manager import CacheManager
monkeypatch.setattr(CacheManager, '_get_writable_cache_dir',
lambda self: str(tmp_path))
monkeypatch.setattr(cache_module, '_retired_writers_warned', set())
cache = CacheManager()
try:
plugin = _plugin('birdnet-go', _manager())
plugin.cache_manager = cache
caplog.set_level(logging.WARNING)
for _ in range(2):
if plugin.request_on_demand(mode='m') is None:
plugin.cache_manager.set('display_on_demand_request', {
'request_id': 'r', 'action': 'start', 'plugin_id': 'birdnet-go'})
assert cache.get('display_on_demand_request', max_age=None, memory_ttl=0) is None
lines = [r.getMessage() for r in caplog.records if 'retired' in r.getMessage()]
assert len(lines) == 1 and "plugin 'birdnet-go'" in lines[0]
finally:
cache.stop_cleanup_thread()
def test_the_feature_detection_pattern(self):
"""The hasattr pattern from docs/PLUGIN_API_REFERENCE.md."""
writes = []
class OldCorePlugin: # an older core's BasePlugin has no such method
pass
for plugin, expect_mailbox in ((OldCorePlugin(), True),
(_plugin('p', _manager()), True),
(_plugin('p', _manager(lambda r: True)), False)):
writes.clear()
if not (hasattr(plugin, 'request_on_demand')
and plugin.request_on_demand(mode='m')):
writes.append('mailbox')
assert (writes == ['mailbox']) is expect_mailbox
class TestArguments:
@@ -389,7 +401,7 @@ class TestArguments:
class TestMockManagers:
def test_a_magicmock_manager_reads_as_not_taken(self):
"""A plugin's test with a MagicMock manager reads as "not taken"."""
"""A plugin's test with a MagicMock manager keeps its mailbox path."""
plugin = _plugin('p', MagicMock())
assert plugin.request_on_demand(mode='m') is None
assert plugin.end_on_demand() is None
+19 -10
View File
@@ -12,8 +12,8 @@ frame, so a command is applied:
* at the next frame in Vegas (8 ms here, at 125 fps).
Each test runs the real run() loop on the fake clock of
test/_run_loop_harness.py. The latencies are the fake clock's, so they are
exact. (The file mailbox these were once compared with is gone: stage 5.)
test/_run_loop_harness.py and compares the socket with the mailbox for the
same request. The latencies are the fake clock's, so they are exact.
"""
import os
@@ -32,10 +32,13 @@ def _event_time(trace, kind):
return next(e[0] for e in trace["events"] if e[1] == kind)
def _on_demand_latency(tmp_path, build, horizon=40):
def _on_demand_latency(tmp_path, build, via, horizon=40):
h = RunLoopHarness(tmp_path, horizon=horizon)
build(h)
h.control_socket().post(POSTED, Command.ON_DEMAND_START, {"plugin_id": "weather"})
if via == "socket":
h.control_socket().post(POSTED, Command.ON_DEMAND_START, {"plugin_id": "weather"})
else:
h.on_demand_request(POSTED, "mb1", plugin_id="weather")
trace = h.run()
return round(_event_time(trace, "on-demand-start") - POSTED, 3), trace
@@ -57,12 +60,18 @@ def _vegas(h):
h.enable_vegas(cycle=30)
@pytest.mark.parametrize("build", [_static, _dwell, _vegas],
ids=["static-screen", "dwell", "vegas"])
def test_a_socket_command_lands_at_once(tmp_path, build):
# Before stage 2 these waited for the next 1 s frame, the next 0.25 s
# dwell tick, or the next 10-frame Vegas check.
socket_latency, trace = _on_demand_latency(tmp_path, build)
@pytest.mark.parametrize("build, mailbox_latency", [
(_static, 0.7), # the next 1 s frame, at t=11
(_dwell, 0.2), # the next 0.25 s tick, at t=10.5
(_vegas, 0.02), # the next 10-frame check: 80 ms at 125 fps, ~0.4 s on a Pi 4
], ids=["static-screen", "dwell", "vegas"])
def test_a_socket_command_lands_at_once(tmp_path, build, mailbox_latency):
(tmp_path / "s").mkdir()
(tmp_path / "m").mkdir()
socket_latency, trace = _on_demand_latency(tmp_path / "s", build, "socket")
via_mailbox, _ = _on_demand_latency(tmp_path / "m", build, "mailbox")
assert via_mailbox == pytest.approx(mailbox_latency, abs=0.002)
if build is _vegas:
# One 8 ms frame: Vegas checks the queue every frame now.
assert socket_latency <= 0.008
+8 -11
View File
@@ -528,17 +528,14 @@ class TestDisplayAPI:
from web_interface.blueprints.api_v3 import api_v3
mock_cache_manager = api_v3.cache_manager = MagicMock()
with patch('web_interface.blueprints.api_v3.display.control_client.on_demand_stop',
side_effect=lambda request_id, **kw: {'accepted': True}) as stop:
response = client.post('/api/v3/display/on-demand/stop')
assert response.status_code == 200
assert response.get_json()['data']['transport'] == 'socket'
stop.assert_called_once()
# The request goes over the control socket; nothing is written to
# the cache (the file mailbox is gone).
mock_cache_manager.set.assert_not_called()
response = client.post('/api/v3/display/on-demand/stop')
# May return 200 if successful or 500 on error
assert response.status_code in [200, 500]
# Verify stop request was set in cache if successful
if response.status_code == 200:
assert mock_cache_manager.set.called
class TestPluginsAPI:
+31 -27
View File
@@ -11,26 +11,32 @@ than even logging it.
import pytest
from src.web_interface.error_handler import describe_exception
from src.web_interface.error_handler import describe_exception, redact_text
class TestDescribeException:
def test_names_the_type_and_message(self):
detail = describe_exception(OSError(5, "Input/output error", "systemctl"))
assert detail == "OSError: [Errno 5] Input/output error: 'systemctl'"
"""describe_exception is a reason code: type and errno, never the message.
def test_the_reported_failure_is_legible(self):
# The whole point: this string is the diagnosis.
assert "Input/output error" in describe_exception(
OSError(5, "Input/output error", "systemctl"))
The message can quote paths, URLs or credentials (CodeQL
py/stack-trace-exposure), so it goes to the log; the code still names the
fault, as "[Errno 5]" did.
"""
def test_an_oserror_names_its_errno(self):
assert describe_exception(
OSError(5, "Input/output error", "systemctl")) == "OSError:EIO"
def test_a_bare_exception_still_names_its_type(self):
# A PermissionError with no message still says more than "unknown".
assert describe_exception(PermissionError()) == "PermissionError"
assert describe_exception(Exception()) == "Exception"
def test_message_is_kept_when_present(self):
assert describe_exception(ValueError("bad port")) == "ValueError: bad port"
def test_the_message_never_reaches_the_code(self):
assert describe_exception(ValueError("bad port /etc/secret")) == "ValueError"
def test_the_message_is_logged_instead(self, caplog):
describe_exception(RuntimeError("disk on fire token=abc123"))
assert "disk on fire" in caplog.text
assert "abc123" not in caplog.text
class TestCredentialRedaction:
@@ -56,47 +62,44 @@ class TestCredentialRedaction:
("authorization: barecredential", "barecredential"),
])
def test_credentials_never_reach_the_response(self, secret_text, leaked):
detail = describe_exception(RuntimeError(secret_text))
detail = redact_text(secret_text)
assert leaked not in detail
assert "<redacted>" in detail
def test_the_parameter_name_survives_redaction(self):
# Knowing *which* credential was involved is part of the diagnosis.
detail = describe_exception(RuntimeError("https://x/y?api_key=SEC123"))
detail = redact_text("https://x/y?api_key=SEC123")
assert "api_key" in detail
def test_unknown_schemes_keep_their_name(self):
for scheme in ("ApiKey", "Negotiate", "NTLM", "AWS4-HMAC-SHA256"):
detail = describe_exception(
RuntimeError("Authorization: %s SECRETVALUE" % scheme))
detail = redact_text("Authorization: %s SECRETVALUE" % scheme)
assert scheme in detail, detail
assert "SECRETVALUE" not in detail, detail
def test_auth_scheme_and_username_survive(self):
# Which kind of credential, and whose, without the credential itself.
assert "Bearer" in describe_exception(
RuntimeError("Authorization: Bearer eyJ.SECRET.sig"))
assert "user" in describe_exception(
RuntimeError("https://user:hunter2@example.com"))
assert "Bearer" in redact_text("Authorization: Bearer eyJ.SECRET.sig")
assert "user" in redact_text("https://user:hunter2@example.com")
def test_non_secret_context_is_preserved(self):
detail = describe_exception(RuntimeError("https://api.x.com/v1?city=Tampa"))
detail = redact_text("https://api.x.com/v1?city=Tampa")
assert "city=Tampa" in detail
assert "<redacted>" not in detail
class TestBounds:
def test_long_messages_are_truncated(self):
detail = describe_exception(ValueError("x" * 5000))
detail = redact_text("x" * 5000)
assert len(detail) <= 400
def test_newlines_are_collapsed_to_one_line(self):
detail = describe_exception(ValueError("line one\nline two\tthree"))
detail = redact_text("line one\nline two\tthree")
assert "\n" not in detail and "\t" not in detail
assert detail == "ValueError: line one line two three"
assert detail == "line one line two three"
def test_custom_length_is_honoured(self):
assert len(describe_exception(ValueError("y" * 500), max_length=50)) <= 50
assert len(redact_text("y" * 500, max_length=50)) <= 50
class TestHandlersCarryDetail:
@@ -319,10 +322,10 @@ class TestHandlersCarryDetail:
assert resp.status_code == 405, "a wrong method must stay a 405"
assert resp.get_json()["error_code"] == "METHOD_NOT_ALLOWED"
# A genuine server fault still reports as one, with its detail.
# A genuine server fault still reports as one, with its reason code.
resp = client.get("/boom")
assert resp.status_code == 500
assert "Input/output error" in resp.get_json()["details"]
assert resp.get_json()["details"] == "OSError:EIO"
def test_global_handler_reports_the_underlying_error(self):
from flask import Flask, jsonify
@@ -345,4 +348,5 @@ class TestHandlersCarryDetail:
client = app.test_client()
body = client.get("/boom").get_json()
assert body["error_code"] == "UNKNOWN_ERROR"
assert "Input/output error" in body["details"]
assert body["details"] == "OSError:EIO"
assert "Input/output error" not in str(body)
@@ -81,6 +81,7 @@ REMOVED_CATCH_ALLS = [
("GET", "/api/v3/config/main", None),
("GET", "/api/v3/config/secrets", None),
("GET", "/api/v3/display/modes", None),
("POST", "/api/v3/display/on-demand/stop", {}),
("GET", "/api/v3/cache/list", None),
("GET", "/api/v3/plugins/installed", None),
("GET", "/api/v3/plugins/health", None),
@@ -113,12 +114,11 @@ def test_the_answer_is_what_the_catch_all_returned(client, caplog, method, url,
assert records[-1].exc_info[1] is FORCED
def test_credentials_are_redacted_from_the_detail(client):
def test_the_exception_message_never_reaches_the_detail(client):
body = client.get("/api/v3/plugins/installed").get_json()
for secret in ("SECRET123", "pw1", "K1"):
assert secret not in body["details"]
assert "<redacted>" in body["details"]
assert body["details"].startswith("RuntimeError: forced failure")
for secret in ("SECRET123", "pw1", "K1", "forced failure"):
assert secret not in str(body)
assert body["details"] == "RuntimeError"
def _raise_415():
@@ -217,7 +217,7 @@ class TestPluginActionStep1:
encoding="utf-8")
return d
def test_the_script_error_reaches_the_response(self, plugin_dir, monkeypatch):
def test_the_script_error_is_reported_by_type(self, plugin_dir, monkeypatch):
from unittest.mock import MagicMock
manager = MagicMock()
manager.get_plugin_directory.return_value = str(plugin_dir)
@@ -231,5 +231,6 @@ class TestPluginActionStep1:
assert resp.status_code == 500
body = resp.get_json()
assert body["details"] == "RuntimeError: the auth script failed"
assert body["details"] == "RuntimeError"
assert "the auth script failed" not in str(body)
assert body["message"] == 'An error occurred; see logs for details'
@@ -946,21 +946,27 @@ class TestTheStoreReportsWhyItIsEmpty:
class TestACrashCarriesItsDetail:
"""Seventeen Starlark handlers answered 5xx with no detail at all."""
"""Seventeen Starlark handlers answered 5xx with no detail at all.
def test_browse_returns_the_exception_detail(self, client):
The detail is a reason code (the exception type), not the exception's
message, which stays in the log (CodeQL py/stack-trace-exposure).
"""
def test_browse_returns_the_reason_code(self, client):
with patch('web_interface.blueprints.api_v3._get_tronbyte_repository_class',
side_effect=ImportError("No module named 'yaml'")):
body = client.get('/api/v3/starlark/repository/browse').get_json()
assert 'yaml' in body.get('details', ''), body
assert body.get('details') == 'ImportError', body
assert 'yaml' not in str(body), body
def test_status_returns_the_exception_detail(self, client):
def test_status_returns_the_reason_code(self, client):
with patch('web_interface.blueprints.api_v3._get_starlark_plugin',
side_effect=RuntimeError("plugin manager is not attached")):
body = client.get('/api/v3/starlark/status').get_json()
assert 'plugin manager is not attached' in body.get('details', ''), body
assert body.get('details') == 'RuntimeError', body
assert 'plugin manager is not attached' not in str(body), body
class TestTheListingIsNotCappedAtOneThousand:
+26 -21
View File
@@ -222,7 +222,7 @@ def _save_config_atomic(config_manager, config_data, create_backup=True):
config_manager.save_config(config_data)
return True, None
except Exception as e:
return False, str(e)
return False, f"Failed to save configuration ({describe_exception(e)})"
def _coerce_to_bool(value):
"""
Coerce a form value to a proper Python boolean.
@@ -246,7 +246,11 @@ def _coerce_to_bool(value):
return value.lower() in ('true', 'on', '1', 'yes')
return False
def _get_display_service_status():
"""Return status information about the ledmatrix service."""
"""Return status information about the ledmatrix service.
active/returncode only: this goes back in API responses, and systemctl's
output (or an exception's text) is logged rather than returned.
"""
try:
result = subprocess.run(
['systemctl', 'is-active', 'ledmatrix'],
@@ -254,26 +258,18 @@ def _get_display_service_status():
text=True,
timeout=3
)
if result.stderr.strip():
logger.debug('systemctl is-active ledmatrix: %s', result.stderr.strip())
return {
'active': result.stdout.strip() == 'active',
'returncode': result.returncode,
'stdout': result.stdout.strip(),
'stderr': result.stderr.strip()
}
except subprocess.TimeoutExpired:
return {
'active': False,
'returncode': -1,
'stdout': '',
'stderr': 'timeout'
}
except Exception as err:
return {
'active': False,
'returncode': -1,
'stdout': '',
'stderr': str(err)
}
logger.warning('systemctl is-active ledmatrix timed out')
return {'active': False, 'returncode': -1}
except Exception:
logger.warning('Could not query ledmatrix.service status', exc_info=True)
return {'active': False, 'returncode': -1}
def _run_systemctl_command(args):
"""Run a systemctl command safely."""
try:
@@ -295,18 +291,26 @@ def _run_systemctl_command(args):
'stderr': 'timeout'
}
except Exception as err:
logger.warning('%s failed', ' '.join(args), exc_info=True)
return {
'returncode': -1,
'stdout': '',
'stderr': str(err)
'stderr': describe_exception(err)
}
def _public_service_result(result):
"""A _run_systemctl_command result fit for a response: no stdout/stderr."""
if result.get('returncode') != 0:
logger.error('systemctl exited %s: %s', result.get('returncode'),
(result.get('stderr') or '').strip())
return {k: v for k, v in result.items() if k not in ('stdout', 'stderr')}
def _ensure_display_service_running():
"""Ensure the ledmatrix display service is running."""
status = _get_display_service_status()
if status.get('active'):
status['started'] = False
return status
result = _run_systemctl_command(['sudo', 'systemctl', 'start', 'ledmatrix.service'])
result = _public_service_result(
_run_systemctl_command(['sudo', 'systemctl', 'start', 'ledmatrix.service']))
service_status = _get_display_service_status()
result['started'] = result.get('returncode') == 0
result['active'] = service_status.get('active')
@@ -314,7 +318,8 @@ def _ensure_display_service_running():
return result
def _stop_display_service():
"""Stop the ledmatrix display service."""
result = _run_systemctl_command(['sudo', 'systemctl', 'stop', 'ledmatrix.service'])
result = _public_service_result(
_run_systemctl_command(['sudo', 'systemctl', 'stop', 'ledmatrix.service']))
status = _get_display_service_status()
result['active'] = status.get('active')
result['status'] = status
@@ -712,7 +717,7 @@ def _do_transactional_uninstall(plugin_id, preserve_config):
success = api_v3.plugin_store_manager.uninstall_plugin(plugin_id)
except Exception as remove_err:
_rollback()
return False, f"Failed to remove plugin {plugin_id}: {remove_err}"
return False, f"Failed to remove plugin {plugin_id} ({describe_exception(remove_err)})"
if not success:
_rollback()
+135 -207
View File
@@ -9,7 +9,7 @@ from web_interface.blueprints.api_v3 import (
_get_display_service_status, _socket_reason_code, _stop_display_service, api_v3,
jsonify, logger, request, uuid,
)
from web_interface import display_preview, display_state, on_demand_dispatch
from web_interface import display_preview, display_state
import web_interface.blueprints.api_v3 as _pkg
from src.ipc import client as control_client
# Read through the module rather than bound by value: tests patch these
@@ -33,97 +33,87 @@ def _cache_manager():
#: How long a start is sent again to a service that systemd reports running
#: but that has no socket yet (a display still loading its plugins, or one
#: someone else just restarted). A cold start gets the dispatcher's own
#: START_WAIT_SECONDS. Either way the route answers at once (202) and the
#: web process's dispatcher does the waiting.
ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS = 10.0
class _NotDelivered(Exception):
"""The display took an on-demand request over the socket and did not
accept it (``busy``, ``invalid_args``, ...) or did not answer in time."""
def __init__(self, reason):
super().__init__(reason)
self.reason = reason
def _dispatcher():
"""The web process's on-demand dispatcher (web_interface/on_demand_dispatch.py)."""
return on_demand_dispatch.get_dispatcher(_send_on_demand)
def _deliver_on_demand(payload):
"""Hand an on-demand request to the display: control socket, else mailbox.
The socket (src/ipc) answers with an acknowledgement as soon as the
display has the command queued for its render thread. The file mailbox
is written only when the socket could not carry the request at all
(``control_client.should_fall_back``): no socket (the display is stopped
or predates it), a refused connection, or a display too old to know the
command. The display reads it within its mailbox poll interval.
def _pending_start_state():
"""A start the dispatcher is still delivering, one it delivered but the
display has not acted on yet, or one it gave up on: what the status
routes report instead of the display's own state (see _shadows). None
when there is none.
A display that had the request and refused it, or did not answer in
time, raises :class:`_NotDelivered`: a mailbox copy would be refused
the same way, or hide a stuck display behind a "success".
A delivered start reads as ``status: "starting"`` with ``delivered:
true`` for at most DELIVERED_SHOWN_SECONDS: the display acknowledges it
when its socket opens and acts on it seconds later, and until then
publishes its own idle state, which would flash in the UI.
Returns ``(transport, socket_error)``: ``'socket'`` and None, or
``'mailbox'`` and the socket failure's reason code.
"""
dispatcher = on_demand_dispatch.current()
status = dispatcher.status() if dispatcher is not None else None
if status is None:
return None
if status.get('status') == 'delivered':
delivered_at = status.get('last_updated') or 0
if _pkg.time.time() - delivered_at > on_demand_dispatch.DELIVERED_SHOWN_SECONDS:
return None
return dict(status, status='starting', delivered=True, delivered_at=delivered_at)
if status.get('status') not in ('starting', 'error'):
return None
return status
try:
if payload['action'] == 'start':
control_client.on_demand_start(
payload['request_id'], payload.get('plugin_id'), payload.get('mode'),
payload.get('duration'), bool(payload.get('pinned', False)))
else:
control_client.on_demand_stop(payload['request_id'])
return 'socket', None
except control_client.ControlError as e:
reason = _socket_reason_code(e.reason)
if not control_client.should_fall_back(e):
logger.warning("The display did not accept on-demand %s %s (%s)",
payload['action'], payload['request_id'], e)
raise _NotDelivered(reason) from None
if reason in _QUIET_SOCKET_REASONS:
logger.debug("On-demand %s via the mailbox: %s", payload['action'], e)
else:
logger.warning("Control socket did not take on-demand %s (%s); "
"using the mailbox", payload['action'], e)
except Exception: # never let the socket path break the route
logger.exception("Control socket client failed; using the mailbox")
reason = 'internal'
_cache_manager().set('display_on_demand_request', payload)
return 'mailbox', reason
def _display_on_demand_state(snapshot):
"""The display's own on-demand state: (state, source). From the socket's
snapshot, else the cache key it also writes; state None when neither."""
state = display_state.on_demand_state(snapshot)
if state is not None:
return state, 'socket'
# memory_ttl=0: the display service writes this key, so only the file
# is current. This process's memory tier would keep serving the first
# copy it read for the full max_age -- "active" for two minutes after
# the display had already stopped.
return _cache_manager().get('display_on_demand_state', max_age=120, memory_ttl=0), 'cache'
def _send_on_demand(payload):
"""Hand an on-demand request to the display over the control socket.
Returns the display's acknowledgement: it has the command queued for
its render thread. Raises ``control_client.ControlError`` when it did
not take it; there is no other way to reach the display (the file
mailbox ``display_on_demand_request`` is gone).
"""
if payload['action'] == 'start':
return control_client.on_demand_start(
payload['request_id'], payload.get('plugin_id'), payload.get('mode'),
payload.get('duration'), bool(payload.get('pinned', False)))
return control_client.on_demand_stop(payload['request_id'])
def _socket_error_response(request_id, action, reason, message=None, **extra):
"""The answer when the display did not take an on-demand request: ``400``
for arguments it refused, else ``503``."""
def _not_delivered_response(request_id, action, reason):
"""The answer when the display had the request and did not accept it."""
status = 400 if reason == 'invalid_args' else 503
data = {'request_id': request_id, 'transport': 'socket', 'socket_error': reason}
data.update(extra)
return jsonify({
'status': 'error',
'message': message or (f'The display service did not accept the on-demand {action} '
f'request ({reason})'),
'data': data,
'message': (f'The display service did not accept the on-demand {action} '
f'request ({reason})'),
'data': {'request_id': request_id, 'transport': 'socket', 'socket_error': reason},
}), status
def _socket_failure_reason(error):
"""The reportable reason code for an exception from _send_on_demand, logged."""
if isinstance(error, control_client.ControlError):
reason = _socket_reason_code(error.reason)
if reason in _QUIET_SOCKET_REASONS:
logger.debug("On-demand request not taken over the control socket: %s", error)
else:
logger.warning("On-demand request not taken over the control socket: %s", error)
return reason
logger.error("Control socket client failed", exc_info=error)
return 'internal'
def _withdraw_on_demand(request_id):
"""Take a start request the route has refused back out of the mailbox.
The display reads the mailbox for an hour without looking at a
request's age, so one left there after an error answer ran whenever the
display next started. Only this request is removed: the mailbox is
re-read and cleared only while it still holds this request_id, as the
display's _consume_on_demand_request does, so a newer request posted in
the meantime stays for the display to take.
"""
cache = _cache_manager()
try:
current = cache.get('display_on_demand_request', max_age=3600, memory_ttl=0)
if isinstance(current, dict) and current.get('request_id') == request_id:
cache.delete('display_on_demand_request')
except Exception: # the route is answering an error already
logger.warning("Could not withdraw on-demand request %s from the mailbox",
request_id, exc_info=True)
@api_v3.route('/display/current', methods=['GET'])
@@ -226,13 +216,16 @@ def get_on_demand_status():
available (``source: "socket"``), else the cache key it also writes
(``source: "cache"``).
"""
state, source = _display_on_demand_state(display_state.read_state())
pending = _pending_start_state()
if pending is not None and _shadows(pending, state):
# A start the web process is still delivering, has delivered but
# the display has not answered yet, or gave up on (start-timeout):
# newer than anything the display has said.
state, source = pending, 'web'
state = display_state.on_demand_state(display_state.read_state())
source = 'socket'
if state is None:
source = 'cache'
cache = _cache_manager()
# memory_ttl=0: the display service writes this key, so only the file
# is current. This process's memory tier would keep serving the first
# copy it read for the full max_age -- "active" for two minutes after
# the display had already stopped.
state = cache.get('display_on_demand_state', max_age=120, memory_ttl=0)
if state is None:
state = {
'active': False,
@@ -248,31 +241,6 @@ def get_on_demand_status():
'source': source,
}
})
def _shadows(pending, state):
"""Whether the web process's start (or its failure) is newer than the
display's on-demand ``state``.
* Still being sent: always.
* Delivered: until the display publishes the state that answers it.
A display that names its request (``request_id``) answers when the
id matches; its startup state, published as the socket opens and so
possibly after the acknowledgement, names no request or an older one
and does not count. An older display without the field answers with
any state published after the delivery.
* Failed: until the display publishes something later.
"""
if not isinstance(state, dict):
return True
if pending.get('status') == 'starting' and not pending.get('delivered'):
return True
if pending.get('delivered') and 'request_id' in state:
return state.get('request_id') != pending.get('request_id')
shown = state.get('last_updated')
if not isinstance(shown, (int, float)) or isinstance(shown, bool):
return True
return shown < (pending.get('last_updated') or 0)
@api_v3.route('/display/on-demand/start', methods=['POST'])
def start_on_demand_display():
"""Request the display controller to run a specific plugin on-demand."""
@@ -321,10 +289,11 @@ def start_on_demand_display():
resolved_plugin,
)
# The request goes over the control socket, the only way to reach the
# display. A display that is not listening (stopped, or still starting)
# is started when start_service asks for it, and the request is sent
# again once its socket is up.
# Deliver the request over the control socket, or post it to the
# mailbox the display process polls (DisplayController.
# _poll_on_demand_requests). Done before any service start: a stopped
# display has no socket, so the request lands in the mailbox, where a
# freshly started display finds it on its first poll.
request_id = data.get('request_id') or str(uuid.uuid4())
request_payload = {
'request_id': request_id,
@@ -335,85 +304,69 @@ def start_on_demand_display():
'pinned': pinned,
'timestamp': _pkg.time.time()
}
# This start supersedes one the dispatcher is still delivering,
# whatever becomes of it: an older request must not land after it.
dispatcher = on_demand_dispatch.current()
if dispatcher is not None:
dispatcher.cancel('superseded')
try:
_send_on_demand(request_payload)
except Exception as e: # pylint: disable=broad-except
error = e
else:
# A socket acknowledgement is the display itself answering: it is
# running and has the request queued, whatever systemd says (a
# display run by hand or in the emulator has no active unit). So
# nothing is checked or started for it. The service is still
# reported the way _ensure_display_service_running reports a
# running one.
transport, socket_error = _deliver_on_demand(request_payload)
except _NotDelivered as e:
return _not_delivered_response(request_id, 'start', e.reason)
# A socket acknowledgement is the display itself answering: it is
# running and has the request queued, whatever systemd says (a display
# run by hand or in the emulator has no active unit). So nothing is
# checked or started for it -- that answered "not running" for a request
# that had already taken effect. The service is still reported the way
# _ensure_display_service_running reports a running one.
if transport == 'socket':
service_result = (dict(_get_display_service_status(), started=False)
if start_service else None)
return _on_demand_started(request_id, resolved_plugin, resolved_mode,
duration, pinned, service_result)
reason = _socket_failure_reason(error)
if not control_client.display_not_listening(error):
# The display had it and refused it, is too old for the command, or
# this web process cannot use the socket at all (switched off, no
# Unix sockets): starting a service would not change that.
return _socket_error_response(request_id, 'start', reason)
duration, pinned, service_result, transport,
socket_error)
service_status = _get_display_service_status()
if not service_status.get('active') and not start_service:
# The request is in the mailbox, and the display reads it whenever
# it next starts: taken back out, or a request answered with this
# error ran later anyway.
_withdraw_on_demand(request_id)
return jsonify({
'status': 'error',
'message': 'Display service is not running. Please start the display service or enable "Start Service" option.',
'service_status': service_status,
'data': {'request_id': request_id, 'transport': 'socket', 'socket_error': reason},
'service_status': service_status
}), 400
# start_service means "start it if it is not running", as the UI's
# checkbox says; _ensure_display_service_running leaves a running
# service alone (restarting it cost seconds of blank panel for nothing).
# Either way the display has no socket yet. The route does not wait for
# it -- a cold start can outlast a client's timeout (the MQTT bridge's is
# 15 s) -- but answers 202 and leaves the sending to the dispatcher,
# whose outcome the status routes report.
wait = ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS
# checkbox says; _ensure_display_service_running leaves a running service
# alone. This used to stop a running service, sleep 1.5s and start it
# again, so every on-demand or "Preview on display" click -- and every
# MQTT on-demand command, which posts here with the default -- cold-
# restarted the display process: every plugin reloaded and the panel was
# blank for seconds. The restart bought nothing. The running process
# looks at this mailbox at least once a second, from its dwell sleep,
# its render loops and Vegas's interrupt check as well as the main loop
# (DisplayController._mailbox_poll_interval), and a restarted one got
# the request the same way: the
# startup path only restores a session the display itself saved
# (display_on_demand_config), so it loaded nothing it would not have had.
service_result = None
if not service_status.get('active'):
if start_service:
service_result = _ensure_display_service_running()
# Check if service actually started
if service_result and not service_result.get('active'):
_withdraw_on_demand(request_id)
return jsonify({
'status': 'error',
'message': 'Failed to start display service. Please check service logs or start it manually.',
'service_result': service_result
}), 500
wait = on_demand_dispatch.START_WAIT_SECONDS
elif start_service:
service_result = dict(service_status, started=False)
_dispatcher().submit(request_payload, wait_seconds=wait)
return jsonify({
'status': 'starting',
'message': ('The display service is starting; the request is sent to it as soon '
'as it is listening. Check the on-demand status for the outcome.'),
'data': {
'request_id': request_id,
'plugin_id': resolved_plugin,
'mode': resolved_mode,
'duration': duration,
'pinned': pinned,
'service': service_result,
'transport': 'socket',
'socket_error': reason,
'pending': True,
'wait_seconds': wait,
},
}), 202
return _on_demand_started(request_id, resolved_plugin, resolved_mode,
duration, pinned, service_result, transport,
socket_error)
def _on_demand_started(request_id, plugin_id, mode, duration, pinned, service_result):
def _on_demand_started(request_id, plugin_id, mode, duration, pinned,
service_result, transport, socket_error):
"""The success answer of /display/on-demand/start."""
response_data = {
'request_id': request_id,
@@ -422,8 +375,10 @@ def _on_demand_started(request_id, plugin_id, mode, duration, pinned, service_re
'duration': duration,
'pinned': pinned,
'service': service_result,
'transport': 'socket',
'transport': transport,
}
if socket_error:
response_data['socket_error'] = socket_error
return jsonify({'status': 'success', 'data': response_data})
@api_v3.route('/display/on-demand/stop', methods=['POST'])
def stop_on_demand_display():
@@ -432,38 +387,23 @@ def stop_on_demand_display():
# _coerce_to_bool: bool("false") is True, which stopped the service.
stop_service = _coerce_to_bool(data.get('stop_service', False))
# The running display takes the stop over the control socket and
# resumes normal rotation in place (_clear_on_demand); nothing is
# restarted.
# The running display takes the stop over the control socket, or reads
# it from the mailbox within ON_DEMAND_POLL_INTERVAL, and resumes normal
# rotation in place (_clear_on_demand); nothing is restarted.
request_id = data.get('request_id') or str(uuid.uuid4())
request_payload = {
'request_id': request_id,
'action': 'stop',
'timestamp': _pkg.time.time()
}
socket_error = None
# A start the web process is still delivering is dropped first: the
# stop is newer, whatever happens to it below.
dispatcher = on_demand_dispatch.current()
cancelled = dispatcher.cancel('requested-stop') if dispatcher is not None else None
try:
_send_on_demand(request_payload)
except Exception as e: # pylint: disable=broad-except
socket_error = _socket_failure_reason(e)
if not stop_service and not (cancelled and control_client.display_not_listening(e)):
if control_client.display_not_listening(e):
service_status = _get_display_service_status()
message = ('Display service is not running, so the stop could not be '
'delivered. If it resumes an on-demand session when it starts, '
'stop it then.' if not service_status.get('active') else
f'The display service is running but its control socket did '
f'not answer ({socket_error}). It may still be starting; '
f'try again shortly.')
return _socket_error_response(request_id, 'stop', socket_error, message,
service=service_status)
return _socket_error_response(request_id, 'stop', socket_error)
transport, socket_error = _deliver_on_demand(request_payload)
except _NotDelivered as e:
if not stop_service:
return _not_delivered_response(request_id, 'stop', e.reason)
# Stopping the service ends on-demand too, whatever the display did
# with the request.
transport, socket_error = 'socket', e.reason
service_result = None
if stop_service:
@@ -472,16 +412,11 @@ def stop_on_demand_display():
response_data = {
'request_id': request_id,
'service': service_result,
'transport': 'socket',
'transport': transport,
}
if cancelled:
# The start it ended never reached the display.
response_data['cancelled_request_id'] = cancelled
if socket_error:
response_data['socket_error'] = socket_error
return jsonify({'status': 'success', 'data': response_data})
@api_v3.route('/display/current-status', methods=['GET'])
def get_current_display_status():
"""Return the display mode/plugin currently intended to be shown.
@@ -511,11 +446,4 @@ def get_current_display_status():
'plugin_id': None,
'last_updated': None,
}
data = dict(state, source=source)
pending = _pending_start_state()
if pending is not None and _shadows(pending, _display_on_demand_state(snapshot)[0]):
# An on-demand start the web process is still delivering, has
# delivered but the display has not acted on yet, or gave up on:
# what the panel is about to show, or why it will not.
data['on_demand_pending'] = pending
return jsonify({'status': 'success', 'data': data})
return jsonify({'status': 'success', 'data': dict(state, source=source)})
+36 -25
View File
@@ -370,13 +370,25 @@ def _redact_error_record(record):
def _read_errors():
return _errors.read_error_report(_errors_cache())
snapshot, clear_request = _errors.read_error_report(_errors_cache())
return snapshot, clear_request
def _send_error_clear(request_id, cutoff):
"""``errors.clear`` over the control socket: the display's answer.
Raises ControlError when it did not take or apply it."""
return _pkg.control_client.errors_clear(request_id, cutoff)
"""``errors.clear`` over the control socket: the display's answer, or None
when the socket could not carry it (no socket, or a display older than
the command) and the clear goes to the mailbox instead. A display that
took the request and failed raises ControlError (no second copy)."""
client = _pkg.control_client
try:
return client.errors_clear(request_id, cutoff)
except client.ControlError as e:
if not client.should_fall_back(e):
raise
_pkg._log_socket_failure('errors.clear', e, _pkg._socket_reason_code(e.reason))
except Exception: # never let the socket path break the route
logger.exception("Control socket client failed clearing errors; using the mailbox")
return None
@api_v3.route('/errors/summary', methods=['GET'])
@@ -390,7 +402,7 @@ def get_error_summary():
until it has reported; ``generated_at`` says when it did.
"""
try:
summary = _errors.error_summary_from_report(_read_errors())
summary = _errors.error_summary_from_report(*_read_errors())
summary['recent_errors'] = [_redact_error_record(r) for r in summary['recent_errors']]
for pattern in summary['active_patterns'].values():
if isinstance(pattern, dict) and isinstance(pattern.get('sample_messages'), list):
@@ -418,7 +430,7 @@ def get_plugin_errors(plugin_id):
recorded errors is "healthy".
"""
try:
health = _errors.plugin_health_from_report(_read_errors(), plugin_id)
health = _errors.plugin_health_from_report(*_read_errors(), plugin_id)
health['last_error'] = _redact_error_record(health['last_error'])
return success_response(data=health, message="Plugin health retrieved")
except Exception as e:
@@ -437,10 +449,9 @@ def clear_old_errors():
max_age_hours: Maximum age in hours (default: 24, max: 8760 = 1 year)
all: true clears every error recorded so far (max_age_hours ignored)
The errors live in the display service, so the clear is sent to it over
the control socket, and it has applied it when this answers. Without a
running display it fails with 503: the errors shown are then the last
run's, and the display's next run starts with none.
The errors live in the display service, so this records a clear request
that it applies within a few seconds. Reads hide the cleared errors from
the moment the request is recorded.
"""
try:
data = request.get_json(silent=True) or {}
@@ -481,28 +492,28 @@ def clear_old_errors():
send=_send_error_clear)
except _pkg.control_client.ControlError as e:
reason = _pkg._socket_reason_code(e.reason)
_pkg._log_socket_failure('errors.clear', e, reason)
if _pkg.control_client.display_not_listening(e):
message = ("The display service is not running (or is still starting), "
"so the errors could not be cleared. The errors shown are "
"from its last run; it starts again with none.")
elif reason in _pkg.control_client.UPGRADE_REASONS:
message = ("The running display service is too old to clear errors; "
"restart it to pick up the installed version")
elif reason in _pkg._QUIET_SOCKET_REASONS:
message = ("The control socket is not available here, so the errors "
"could not be cleared")
else:
message = "The display service did not apply the clear"
logger.warning("The display did not apply the error clear: %s", e)
return error_response(
error_code=ErrorCode.SYSTEM_ERROR,
message=message,
message="The display service did not apply the clear",
context={'socket_error': reason},
status_code=503
)
except OSError as e:
logger.error("Could not record an error clear request: %s", e)
return error_response(
error_code=ErrorCode.SYSTEM_ERROR,
message="Could not record the clear request in the shared cache",
status_code=500
)
scope = "all errors" if clear_all else f"errors older than {max_age_hours} hours"
return success_response(data=result, message=f"Cleared {scope}")
if result.get('applied'):
message = f"Cleared {scope}"
else:
message = (f"Clear of {scope} requested; the display service applies it "
f"within about {int(_errors.SNAPSHOT_TICK_INTERVAL)} seconds")
return success_response(data=result, message=message)
except Exception as e:
logger.error(f"Error clearing old errors: {e}", exc_info=True)
return error_response(
@@ -983,7 +983,6 @@ def stop_pixlet_editor():
'status': 'error',
'message': 'Editor force-stopped, but the display could not be '
'restarted automatically - start it manually.',
'details': (result.get('stderr') or '').strip(),
'data': {'running': False}}), 500
return jsonify({'status': 'success',
'message': 'Editor force-stopped; the display has been '
+3 -4
View File
@@ -689,7 +689,7 @@ def execute_system_action():
logger.warning("install_base_requirements timed out for %s", label)
except OSError as install_err:
all_ok = False
outputs.append(f"== {label} ==\nFailed: {install_err}")
outputs.append(f"== {label} ==\nFailed: {describe_exception(install_err)}")
logger.warning("install_base_requirements errored for %s: %s", label, install_err)
return jsonify({
'status': 'success' if all_ok else 'error',
@@ -784,11 +784,10 @@ def execute_system_action():
return jsonify({'status': 'error', 'message': 'Command timed out', 'returncode': -1, 'stderr': 'timeout'})
except Exception as e:
logger.error("execute_system_action failed: %s", e, exc_info=True)
detail = describe_exception(e)
resp = {
'status': 'error',
'message': _sudo_hint_for(detail) or 'Action failed; see logs for details',
'details': detail,
'message': _sudo_hint_for(str(e)) or 'Action failed; see logs for details',
'details': describe_exception(e),
}
return jsonify(resp), 500
@api_v3.route('/system/git-info', methods=['GET'])
+2 -2
View File
@@ -67,7 +67,7 @@ def _run_background_connect(ssid, password):
payload = _connect_result_payload(ssid, success, message)
except Exception as e:
logger.error("Background WiFi connect failed", exc_info=True)
payload = {'status': 'error', 'message': describe_exception(e)}
payload = {'status': 'error', 'message': f'Failed to connect to network ({describe_exception(e)})'}
_record_connect_result(ssid, payload)
@@ -276,7 +276,7 @@ def connect_wifi():
try:
success, message = wifi_manager.connect_to_network(ssid, password)
except Exception as e:
_record_connect_result(ssid, {'status': 'error', 'message': describe_exception(e)})
_record_connect_result(ssid, {'status': 'error', 'message': f'Failed to connect to network ({describe_exception(e)})'})
raise
payload = _connect_result_payload(ssid, success, message)
_record_connect_result(ssid, payload)
-239
View File
@@ -1,239 +0,0 @@
"""Deliver an on-demand start to a display that is not listening yet.
``POST /api/v3/display/on-demand/start`` can find no display on the control
socket: the service is stopped (the route starts it) or still loading its
plugins. The socket comes up only when the display's run loop starts, which
can take longer than a client waits -- the MQTT bridge gives up after 15 s.
So the route answers at once (``202``, ``status: "starting"``) and hands
the request to the one :class:`OnDemandDispatcher` of the web process, whose
worker thread sends it again until the display acknowledges it or
:data:`START_WAIT_SECONDS` pass.
* One request at a time: a newer start replaces the pending one, and a
stop cancels it (:meth:`OnDemandDispatcher.cancel`).
* Its outcome is :meth:`OnDemandDispatcher.status`, which
``GET /display/on-demand/status`` and ``/display/current-status`` report:
``starting`` while it waits, ``delivered`` once acknowledged (the
display's own state takes over from there), or ``error`` with
``start-timeout`` or the socket's reason.
Nothing is written to disk: the file mailbox that once carried such a
request is gone (docs/IPC_CONTROL_SOCKET.md, stage 5).
"""
from __future__ import annotations
import threading
import time
from typing import Any, Callable, Dict, Optional
from src.ipc import client as control_client
from src.logging_config import get_logger
logger = get_logger(__name__)
#: How long the worker keeps sending a start before it gives up
#: (``start-timeout``). The socket comes up when the display's run loop
#: starts, after every plugin has loaded.
START_WAIT_SECONDS = 45.0
#: Gap between two sends while nothing is listening.
RETRY_INTERVAL = 0.5
#: How long a finished outcome (delivered, error, cancelled) is still
#: reported, so a client polling every few seconds sees it.
OUTCOME_SECONDS = 120.0
#: How long a delivered start is still reported as ``starting`` while the
#: display has not published the state that answers it. The display
#: acknowledges a start when its socket opens, but its run loop acts on it
#: only after its first screen is built (about 5 s on a Pi 4 in Vegas), and
#: until then it reports its own idle state. See
#: ``web_interface/blueprints/api_v3/display.py:_pending_start_state``.
DELIVERED_SHOWN_SECONDS = 30.0
#: ``send(payload)`` hands the request to the display (the route's
#: ``_send_on_demand``) and raises ``ControlError`` when it does not take it.
Sender = Callable[[Dict[str, Any]], Any]
class OnDemandDispatcher:
"""One pending on-demand start, and the worker thread that delivers it."""
def __init__(self, send: Sender, *,
wait_seconds: Optional[float] = None,
retry_interval: Optional[float] = None,
clock: Callable[[], float] = time.monotonic,
wall_clock: Callable[[], float] = time.time):
self._send = send
self.wait_seconds = START_WAIT_SECONDS if wait_seconds is None else wait_seconds
self.retry_interval = RETRY_INTERVAL if retry_interval is None else retry_interval
self._clock = clock
self._wall = wall_clock
self._lock = threading.Lock()
# Held by the worker while it reads the pending start and sends it,
# so cancel() can wait out a send already in flight: a cancelled (or
# superseded) start must not reach the display after the request
# that replaced it. Never taken while holding _lock.
self._send_lock = threading.Lock()
self._wake = threading.Event()
# Bumped by every submit and cancel: a send that started under an
# older generation does not report its result as the current one.
self._generation = 0
self._pending: Optional[Dict[str, Any]] = None
self._deadline = 0.0
self._status: Optional[Dict[str, Any]] = None
self._finished_at: Optional[float] = None
self._thread: Optional[threading.Thread] = None
# -- the routes' side ------------------------------------------------------
def submit(self, payload: Dict[str, Any], wait_seconds: Optional[float] = None) -> None:
"""Deliver ``payload`` (an on-demand start) in the background,
replacing any start still pending."""
with self._lock:
self._generation += 1
superseded = self._pending
self._pending = dict(payload)
self._deadline = self._clock() + (self.wait_seconds if wait_seconds is None
else wait_seconds)
self._status = self._describe('starting', payload)
self._finished_at = None
if self._thread is None or not self._thread.is_alive():
self._thread = threading.Thread(target=self._run, name='on-demand-dispatch',
daemon=True)
self._thread.start()
if superseded is not None:
logger.info("On-demand start %s superseded by %s before the display took it",
superseded.get('request_id'), payload.get('request_id'))
self._wake.set()
def cancel(self, reason: str = 'cancelled') -> Optional[str]:
"""Drop the pending start (a stop arrived). Returns its request id,
or None when nothing was pending."""
with self._lock:
pending = self._pending
if pending is None:
return None
self._generation += 1
self._pending = None
self._finish(self._describe('idle', pending, last_event=reason))
self._wake.set()
# A send of it may be in flight (at most the client's timeout, 1 s):
# let it finish, so whatever the caller sends next lands after it.
with self._send_lock:
pass
logger.info("On-demand start %s cancelled before the display took it (%s)",
pending.get('request_id'), reason)
return pending.get('request_id')
def status(self) -> Optional[Dict[str, Any]]:
"""The pending start's state, or its outcome for OUTCOME_SECONDS
after it finished; None otherwise. In the shape of the display's
on-demand state (``active``, ``status``, ``error``, ...), plus
``source: "web"``."""
with self._lock:
if self._status is None:
return None
if (self._finished_at is not None
and self._clock() - self._finished_at > OUTCOME_SECONDS):
return None
return dict(self._status)
def pending(self) -> bool:
with self._lock:
return self._pending is not None
# -- the worker ------------------------------------------------------------
def _describe(self, status: str, payload: Dict[str, Any], error: Optional[str] = None,
last_event: Optional[str] = None) -> Dict[str, Any]:
return {
'active': False,
'status': status,
'error': error,
'last_event': last_event,
'request_id': payload.get('request_id'),
'plugin_id': payload.get('plugin_id'),
'mode': payload.get('mode'),
'duration': payload.get('duration'),
'pinned': bool(payload.get('pinned', False)),
'last_updated': self._wall(),
'source': 'web',
}
def _finish(self, status: Dict[str, Any]) -> None:
"""Record an outcome. Caller holds _lock."""
self._status = status
self._finished_at = self._clock()
def _run(self) -> None:
while True:
with self._send_lock:
with self._lock:
payload, generation = self._pending, self._generation
deadline = self._deadline
if payload is None:
self._thread = None
return
outcome, error = self._attempt(payload)
with self._lock:
if generation != self._generation:
continue # superseded or cancelled meanwhile
if outcome == 'retry' and self._clock() + self.retry_interval > deadline:
outcome, error = 'error', 'start-timeout'
if outcome == 'delivered':
self._pending = None
self._finish(self._describe('delivered', payload,
last_event='delivered'))
elif outcome == 'error':
self._pending = None
self._finish(self._describe('error', payload, error=error))
else:
self._wake.clear()
if outcome == 'delivered':
logger.info("On-demand start %s delivered once the display was listening",
payload.get('request_id'))
elif outcome == 'error':
logger.warning("On-demand start %s not delivered: %s",
payload.get('request_id'), error)
else:
# A submit or cancel wakes the wait at once.
self._wake.wait(self.retry_interval)
def _attempt(self, payload: Dict[str, Any]):
try:
self._send(payload)
except control_client.ControlError as e:
if control_client.display_not_listening(e):
return 'retry', None
return 'error', str(e.reason)
except Exception: # pylint: disable=broad-except
logger.exception("On-demand start %s: the control socket client failed",
payload.get('request_id'))
return 'error', 'internal'
return 'delivered', None
_dispatcher: Optional[OnDemandDispatcher] = None
_dispatcher_lock = threading.Lock()
def get_dispatcher(send: Sender) -> OnDemandDispatcher:
"""The web process's dispatcher, created on first use."""
global _dispatcher
with _dispatcher_lock:
if _dispatcher is None:
_dispatcher = OnDemandDispatcher(send)
return _dispatcher
def current() -> Optional[OnDemandDispatcher]:
"""The dispatcher if one was created, without creating one."""
return _dispatcher
def reset_for_tests() -> None:
global _dispatcher
with _dispatcher_lock:
_dispatcher = None
+2 -5
View File
@@ -356,12 +356,9 @@ window.previewPluginNow = function(pluginId) {
})
.then(r => r.json())
.then(data => {
// 'starting' (202): the display service is starting and the request
// follows once it listens.
const starting = data.status === 'starting';
showNotification(data.message || ('Previewing ' + pluginId + ' for 60 seconds'),
starting ? 'info' : (data.status || 'success'));
if (data.status === 'success' || starting) window.toggleFloatingPreview(true);
data.status || 'success');
if (data.status === 'success') window.toggleFloatingPreview(true);
})
.catch(err => {
showNotification('Preview failed: ' + err.message, 'error');
@@ -1739,13 +1739,6 @@ function submitOnDemandRequest(event) {
showNotification(`Requested on-demand mode for ${pluginName}`, 'success');
closeOnDemandModal();
setTimeout(() => loadOnDemandStatus(true), 700);
} else if (result.status === 'starting') {
// 202: the display service is starting; the web process sends
// the request once it listens. The status card shows the outcome.
const pluginName = resolvePluginDisplayName(currentOnDemandPluginId);
showNotification(`Starting the display for ${pluginName}…`, 'info');
closeOnDemandModal();
setTimeout(() => loadOnDemandStatus(true), 700);
} else {
console.error('[submitOnDemandRequest] Request failed:', result);
showNotification(result.message || 'Failed to start on-demand mode', 'error');
+2 -2
View File
@@ -86,7 +86,7 @@ def refresh_after_update(run=None, systemd_dir=None, helper_source=None, helper_
except Exception as e: # a broken template must not fail the update itself
if type(e).__name__ != 'UnitsUnreadable':
logger.warning("Could not compare the installed systemd units with the new templates: %s", e)
return _result(FAILED, f'The service settings could not be checked: {e}.')
return _result(FAILED, 'The service settings could not be checked; see logs for details.')
# Units installed mode 0600 (install_service.sh run on its own, before
# it set 0644): only root can compare them, so let the helper decide.
stale = []
@@ -110,7 +110,7 @@ def refresh_after_update(run=None, systemd_dir=None, helper_source=None, helper_
timeout=TIMEOUT_SECONDS)
except (subprocess.SubprocessError, OSError) as e:
logger.warning("Refreshing the systemd units failed: %s", e)
return _result(FAILED, f'Updating the service settings ({names}) failed: {e}.', stale)
return _result(FAILED, f'Updating the service settings ({names}) failed; see logs for details.', stale)
if result.returncode == 0:
# The helper says what it did: "units refreshed: a b" or "units: up to date".
done = next((line.split(':', 1)[1].split() for line in (result.stdout or '').splitlines()