mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-05 14:55:08 +00:00
feat(web): a start that waits for the display answers 202 and is delivered in the background
The start route held a request open for up to 45 s while a cold-started display loaded its plugins; the MQTT bridge (15 s timeout) and browsers reported a failure for a request that was then delivered. Now, when no display is listening, the route starts the service if asked and answers 202 with status "starting" at once. A single worker in the web process (web_interface/on_demand_dispatch.py) sends the request until the display acknowledges it or the wait runs out (45 s cold start, 10 s for a running service without a socket yet). A newer start supersedes the pending one; a stop cancels it (and succeeds, with cancelled_request_id, even with no display listening). The outcome is reported by /display/on-demand/status (source "web": starting, or error with start-timeout or the socket's reason, until the display publishes something newer) and by /display/current-status as on_demand_pending. Callers: the web UI's on-demand modal and "Preview on display" treat "starting" as taken (an info toast); the MQTT bridge already treats any non-error 2xx as success (now pinned by a test). Tests: the dispatcher (ack, retry then ack, start-timeout, other failures, superseded, an in-flight ack for a superseded start, stop while pending, a per-start wait, outcome lifetime); the routes (202, status routes while pending and after a timeout, a later display state replacing the failure, stop while pending, a new start superseding); a JS suite for app.js. Mutation check: 20 mutants on the worker, the routes and app.js, 20 killed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
+11
-5
@@ -42,11 +42,17 @@ to for one release are gone.
|
||||
pomodoro-timer still write the mailbox, but only as their fallback when
|
||||
those methods are missing or answer `None`.
|
||||
- **On-demand routes without a listening display.** `POST
|
||||
/api/v3/display/on-demand/start` with the service stopped starts it (when
|
||||
`start_service`, the default) and sends the request once the display's
|
||||
socket answers, waiting up to 45 s (10 s for a service that is running but
|
||||
has no socket yet); otherwise it answers `400` (`start_service` false) or
|
||||
`503` with `socket_error`. Every other socket failure (`unknown_command`
|
||||
/api/v3/display/on-demand/start` with no display listening starts the
|
||||
service (when `start_service`, the default) and answers **`202`** with
|
||||
`status: "starting"` at once; a single background worker in the web
|
||||
process (`web_interface/on_demand_dispatch.py`) sends the request until
|
||||
the display acknowledges it, for up to 45 s (10 s for a service that is
|
||||
running but has no socket yet). `GET /display/on-demand/status` reports it
|
||||
(`starting`, then the display's state, or `error` / `start-timeout`), and
|
||||
`/display/current-status` adds `on_demand_pending`. A newer start replaces
|
||||
a pending one and a stop cancels it (`cancelled_request_id`). The web UI
|
||||
and the MQTT bridge treat `202` as taken. With the service stopped and
|
||||
`start_service` false it answers `400`. Every other socket failure (`unknown_command`
|
||||
from an older display, `disabled`/`unsupported`, `busy`, a timeout) is a
|
||||
`503`. `/stop` answers `503` when no display is listening, unless
|
||||
`stop_service` stops the service. `transport` is always `"socket"`; the
|
||||
|
||||
@@ -62,8 +62,10 @@ send the command over the display's control socket and get an ack. That is
|
||||
the only way in: the cache-key mailboxes (`display_on_demand_request`,
|
||||
`plugin_error_clear_request`) are gone, and a write to either is dropped
|
||||
with a warning. When no display is listening yet, the start route starts the
|
||||
service (if asked) and sends the request again once the socket is up; any
|
||||
other failure is answered as an error. Socket commands and plugins'
|
||||
service (if asked) and answers `202`; the web process's dispatcher
|
||||
([`on_demand_dispatch.py`](../web_interface/on_demand_dispatch.py)) sends the
|
||||
request once the socket is up, and the status routes report the outcome.
|
||||
Any other failure is answered as an error. Socket commands and plugins'
|
||||
in-process requests end in the same handler, `_handle_on_demand_request()`.
|
||||
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
|
||||
for the protocol and the permission model.
|
||||
|
||||
+27
-10
@@ -472,19 +472,35 @@ request was never sent.
|
||||
|
||||
| What happened | Example reasons | On-demand start | On-demand stop | `errors.clear` |
|
||||
|---|---|---|---|---|
|
||||
| No display listening | `no_socket`, `refused` | service stopped: `400` without `start_service`; with it, start the service and send again once the socket answers (up to 45 s), else `503`. Service running (still starting): send again for up to 10 s, else `503` | `503` ("not running" / "may still be starting"); with `stop_service` the service is stopped and the route succeeds | `503` ("not running"; its errors are the last run's, and the next run starts with none) |
|
||||
| No display listening | `no_socket`, `refused` | service stopped: `400` without `start_service`; with it, start the service and answer `202` (`status: "starting"`) at once; the dispatcher sends the request until the display takes it (up to 45 s), else `start-timeout`. Service running (still starting): the same `202`, sent for up to 10 s | `503` ("not running" / "may still be starting"); with `stop_service` the service is stopped and the route succeeds | `503` ("not running"; its errors are the last run's, and the next run starts with none) |
|
||||
| A display too old to know the command | `unknown_command`, `unsupported_version` | `503` | `503` | `503`, "restart it" |
|
||||
| No socket in this process | `disabled`, `unsupported` (Windows, `LEDMATRIX_CONTROL_SOCKET=off`) | `503` | `503` | `503` |
|
||||
| The display had it and failed, turned it away, or never answered | `busy`, `invalid_args`, `internal`, a timeout, a hang-up, `bad_response`, `forbidden` | `503` (`400` for `invalid_args`) | `503` (with `stop_service`: stopped anyway) | `503` |
|
||||
|
||||
Every error answer carries `socket_error` (a reason code, or `other`).
|
||||
Nothing is written to the cache in any of these cases. The start route's
|
||||
waits are bounded (`ON_DEMAND_SOCKET_WAIT_SECONDS`,
|
||||
`ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS` in
|
||||
`web_interface/blueprints/api_v3/display.py`): the socket comes up when the
|
||||
display's run loop starts, after every plugin has loaded. A client with a
|
||||
shorter HTTP timeout (the MQTT bridge's is 15 s) can give up first while
|
||||
the route still delivers the request.
|
||||
Nothing is written to the cache in any of these cases.
|
||||
|
||||
**Waiting for a display that is starting.** The socket comes up when the
|
||||
display's run loop starts, after every plugin has loaded, which can take
|
||||
longer than a client waits (the MQTT bridge gives up after 15 s). So the
|
||||
start route never waits: it answers `202` with `status: "starting"`, and
|
||||
hands the request to the web process's one dispatcher
|
||||
([`web_interface/on_demand_dispatch.py`](../web_interface/on_demand_dispatch.py)).
|
||||
Its worker thread sends the request every 0.5 s while nothing is listening,
|
||||
until the display acknowledges it or the wait runs out (45 s after a cold
|
||||
start, `START_WAIT_SECONDS`; 10 s for a service that was already running,
|
||||
`ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS`). Any other failure ends it at once.
|
||||
One start is pending at a time: a newer start replaces it, and a stop
|
||||
cancels it (the stop then succeeds even with no display listening, and
|
||||
reports `cancelled_request_id`).
|
||||
|
||||
The outcome is reported where clients already look:
|
||||
`GET /display/on-demand/status` answers the pending start's state
|
||||
(`source: "web"`, `status: "starting"`, or `status: "error"` with `error:
|
||||
"start-timeout"` or the socket's reason) until the display publishes
|
||||
something newer, and `GET /display/current-status` adds it as
|
||||
`on_demand_pending`. Once the display has taken the request its own state
|
||||
is reported, as for any start.
|
||||
|
||||
Brightness and plugin reload never had a mailbox: without the socket, the
|
||||
config watcher applies the saved brightness and a reload becomes the
|
||||
@@ -676,8 +692,9 @@ device never touches the live display.
|
||||
interface no longer writes `display_on_demand_request` or
|
||||
`plugin_error_clear_request`, and the display no longer reads them (see
|
||||
"Without the socket"). When no display is listening, the start route
|
||||
starts the service if asked and sends the request again once its socket
|
||||
is up; every other failure is an error the route reports. A write to
|
||||
starts the service if asked, answers `202`, and the web process's
|
||||
dispatcher sends the request once the socket is up; every other failure
|
||||
is an error the route reports. A write to
|
||||
either key is dropped with a one-time warning naming the writer. The
|
||||
display still writes `display_current_state`, `display_on_demand_state`
|
||||
and `plugin_runtime_snapshot`: the web interface reads them whenever the
|
||||
|
||||
+32
-10
@@ -464,7 +464,7 @@ Request a specific plugin to display on-demand.
|
||||
- `mode` (string, optional): Display mode name (plugin_id inferred if not provided)
|
||||
- `duration` (number, optional): Duration in seconds (0 = until stopped)
|
||||
- `pinned` (boolean, optional): Pin display (pause rotation)
|
||||
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket. A stopped one is started and sent the request once its socket is up, which can take as long as the display takes to load its plugins (the route waits up to 45 s). When false and the service is stopped, the route returns 400.
|
||||
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket. A stopped one is started, and the route answers `202` at once (see below); the request is sent once the display's socket is up, which can take as long as the display takes to load its plugins. When false and the service is stopped, the route returns 400.
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
@@ -490,15 +490,35 @@ applies it within a frame; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)).
|
||||
The `"mailbox"` value earlier releases could answer is gone with the
|
||||
mailbox: nothing is written to the cache.
|
||||
|
||||
When the display did not take the request, the route answers an error with
|
||||
`status: "error"` and `data: {request_id, transport: "socket",
|
||||
socket_error}` (plus `service` when it started or checked the service):
|
||||
**No display listening yet** (`no_socket`, `refused`: the service is
|
||||
stopped, or still loading its plugins). With the service stopped and
|
||||
`start_service` false, `400`. Otherwise the route starts the service if
|
||||
needed and answers at once:
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "starting",
|
||||
"message": "The display service is starting; ...",
|
||||
"data": {"request_id": "uuid-here", "plugin_id": "football-scoreboard", "mode": "nfl_live",
|
||||
"duration": 45, "pinned": true, "service": {"active": true, "started": true},
|
||||
"transport": "socket", "socket_error": "no_socket",
|
||||
"pending": true, "wait_seconds": 45.0}
|
||||
}
|
||||
```
|
||||
|
||||
with HTTP `202`. The web process sends the request until the display takes
|
||||
it, for up to `wait_seconds` (45 after a cold start, 10 when the service was
|
||||
already running). Follow it with `GET /api/v3/display/on-demand/status`:
|
||||
its `state` is `{status: "starting", source: "web", request_id, ...}` while
|
||||
it waits, the display's own state once delivered, or `{status: "error",
|
||||
error: "start-timeout"}` (or the socket's reason) if it never was;
|
||||
`GET /api/v3/display/current-status` carries the same as
|
||||
`on_demand_pending`. A newer start replaces a pending one; a stop cancels it.
|
||||
|
||||
Otherwise, when the display did not take the request, the route answers an
|
||||
error with `status: "error"` and `data: {request_id, transport: "socket",
|
||||
socket_error}`:
|
||||
|
||||
- no display listening (`no_socket`, `refused`): with the service stopped
|
||||
and `start_service` false, `400`; otherwise the route waits for the
|
||||
display's socket (45 s after starting the service, 10 s when it was
|
||||
already running and may still be starting) and answers `503` if it never
|
||||
answers;
|
||||
- a full queue (`busy`), no answer after the request was sent (`timeout`,
|
||||
`closed`), a display older than the command (`unknown_command`), no
|
||||
socket in the web process (`disabled`, `unsupported`): `503` at once;
|
||||
@@ -507,7 +527,9 @@ socket_error}` (plus `service` when it started or checked the service):
|
||||
The stop route answers the same errors (`503` when no display is
|
||||
listening, with a message saying whether the service is stopped), except
|
||||
that with `stop_service: true` it still stops the service and answers
|
||||
success, with `socket_error` set.
|
||||
success, with `socket_error` set, and that a stop which cancelled a pending
|
||||
start succeeds with `cancelled_request_id` even when no display is
|
||||
listening.
|
||||
|
||||
### Stop On-Demand Display
|
||||
|
||||
|
||||
@@ -328,6 +328,20 @@ def _hermetic_control_socket(monkeypatch):
|
||||
monkeypatch.setenv(SOCKET_PATH_ENV, 'off')
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _no_pending_on_demand_dispatch():
|
||||
"""Drop the web process's on-demand dispatcher after each test, so a
|
||||
start one test left pending is not still being sent in the next."""
|
||||
yield
|
||||
module = sys.modules.get('web_interface.on_demand_dispatch')
|
||||
if module is None:
|
||||
return
|
||||
dispatcher = module.current()
|
||||
if dispatcher is not None:
|
||||
dispatcher.cancel('test-teardown')
|
||||
module.reset_for_tests()
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _hermetic_unit_refresh(monkeypatch, tmp_path_factory):
|
||||
"""Keep updates' systemd unit refresh off the host.
|
||||
|
||||
+2
-1
@@ -32,7 +32,8 @@ const UNIT = ['unit/test_list_filter.js', 'unit/test_render_cards.js',
|
||||
'unit/test_page_registry.js', 'unit/test_core_modules.js',
|
||||
'unit/test_overview_reconciliation_poll.js',
|
||||
'unit/test_display_partial_ids.js',
|
||||
'unit/test_general_web_login_token.js'];
|
||||
'unit/test_general_web_login_token.js',
|
||||
'unit/test_on_demand_starting.js'];
|
||||
const DOM = ['dom/test_installed_dom.js', 'dom/test_store_dom.js', 'dom/test_no_double_fetch.js',
|
||||
'dom/test_tools_sections.js', 'dom/test_cache_page.js',
|
||||
'dom/test_durations_page.js', 'dom/test_operation_history_page.js',
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
// POST /api/v3/display/on-demand/start answers 202 with status "starting"
|
||||
// when the display service has to be started first: the request is taken,
|
||||
// and the web process sends it once the display listens. "Preview on
|
||||
// display" (app.js) must read that as taken -- an info toast and the
|
||||
// floating preview opened -- not as a failure. Runs the shipped app.js in a
|
||||
// vm with a minimal fake DOM, as test_restart_banner.js does.
|
||||
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const vm = require('vm');
|
||||
const V3 = path.resolve(__dirname, '../../../web_interface/static/v3');
|
||||
|
||||
let pass = 0, fail = 0;
|
||||
const ok = (label, cond, extra) => cond
|
||||
? (pass++, console.log(' ok ' + label))
|
||||
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' ' + JSON.stringify(extra) : '')));
|
||||
|
||||
function load(answer) {
|
||||
const notes = [];
|
||||
const opened = [];
|
||||
const noop = () => {};
|
||||
const document = {
|
||||
body: { addEventListener: noop },
|
||||
addEventListener: noop,
|
||||
getElementById: () => null,
|
||||
querySelector: () => null,
|
||||
querySelectorAll: () => [],
|
||||
};
|
||||
const window = { addEventListener: noop, getApp: () => null };
|
||||
const context = {
|
||||
window, document, console,
|
||||
sessionStorage: { setItem: noop, removeItem: noop, getItem: () => null },
|
||||
showNotification: (m, t) => notes.push([m, t]),
|
||||
setTimeout: () => 0,
|
||||
fetch: () => Promise.resolve({ json: () => Promise.resolve(answer) }),
|
||||
};
|
||||
vm.createContext(context);
|
||||
vm.runInContext(fs.readFileSync(path.join(V3, 'app.js'), 'utf8'), context);
|
||||
window.toggleFloatingPreview = (open) => opened.push(open);
|
||||
return { window, notes, opened };
|
||||
}
|
||||
|
||||
async function preview(answer) {
|
||||
const t = load(answer);
|
||||
t.window.previewPluginNow('weather');
|
||||
for (let i = 0; i < 5; i++) await Promise.resolve();
|
||||
return t;
|
||||
}
|
||||
|
||||
(async () => {
|
||||
console.log('\npreviewPluginNow');
|
||||
{
|
||||
const t = await preview({ status: 'starting', message: 'The display service is starting',
|
||||
data: { request_id: 'r1', pending: true } });
|
||||
ok('a 202 "starting" answer is an info toast, not an error',
|
||||
t.notes.length === 1 && t.notes[0][1] === 'info', t.notes);
|
||||
ok('and the preview opens', t.opened.length === 1 && t.opened[0] === true, t.opened);
|
||||
}
|
||||
{
|
||||
const t = await preview({ status: 'success', data: { request_id: 'r1' } });
|
||||
ok('a 200 success still opens it', t.opened.length === 1 && t.notes[0][1] === 'success', t.notes);
|
||||
}
|
||||
{
|
||||
const t = await preview({ status: 'error', message: 'no display' });
|
||||
ok('an error does not', t.opened.length === 0 && t.notes[0][1] === 'error', t.notes);
|
||||
}
|
||||
console.log(`\n${pass} passed, ${fail} failed`);
|
||||
process.exit(fail ? 1 : 0);
|
||||
})();
|
||||
@@ -16,7 +16,8 @@ This file previously pinned that restart path (it guarded a broken
|
||||
``import _pkg.time`` inside it). The path is gone; these tests pin its
|
||||
replacement: a running service is left alone, a stopped one is started (only
|
||||
when start_service is set), and the request goes over the control socket
|
||||
either way -- to a stopped display once it has started and its socket is up.
|
||||
either way -- to a stopped display once it has started and its socket is up,
|
||||
sent by the web process's dispatcher after the route has answered 202.
|
||||
Nothing is ever written to the cache: the file mailbox is gone (stage 5).
|
||||
|
||||
The service helpers are patched where they run. display.py binds
|
||||
@@ -26,6 +27,7 @@ _run_systemctl_command is the one place a systemctl command is issued.
|
||||
"""
|
||||
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
@@ -157,10 +159,18 @@ class TestStartWhileTheServiceIsStopped:
|
||||
def test_start_service_starts_it_once_and_never_stops_it(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
response = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
assert response.status_code == 200, response.get_json()
|
||||
# Answered at once: the web process sends the request in the
|
||||
# background once the started display listens.
|
||||
assert response.status_code == 202, response.get_json()
|
||||
assert response.get_json()["status"] == "starting"
|
||||
assert _systemctl_verbs(service["systemctl"]) == ["start"]
|
||||
service["stop_service"].assert_not_called()
|
||||
# Sent once the started display's socket answered.
|
||||
from web_interface import on_demand_dispatch
|
||||
dispatcher = on_demand_dispatch.current()
|
||||
deadline = time.monotonic() + 5
|
||||
while dispatcher.pending() and time.monotonic() < deadline:
|
||||
time.sleep(0.01)
|
||||
assert dispatcher.status()["status"] == "delivered"
|
||||
assert [s[0] for s in service["sent"]] == ["start"]
|
||||
assert _mailbox_writes(service["cache"]) == []
|
||||
|
||||
|
||||
@@ -14,6 +14,7 @@ runs a real server on a temp socket (Linux/macOS only).
|
||||
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
@@ -139,89 +140,175 @@ def _attempts(*outcomes, calls=None, clock=None):
|
||||
return attempt
|
||||
|
||||
|
||||
def _until(predicate, timeout=5.0):
|
||||
end = time.monotonic() + timeout
|
||||
while time.monotonic() < end:
|
||||
if predicate():
|
||||
return True
|
||||
time.sleep(0.005)
|
||||
return False
|
||||
|
||||
|
||||
def _no_socket():
|
||||
return control_client.ControlError("no_socket", "x", sent=False)
|
||||
|
||||
|
||||
class TestNoDisplayListening:
|
||||
"""No socket to talk to: the display is stopped, or still starting."""
|
||||
"""No socket to talk to: the display is stopped, or still starting.
|
||||
|
||||
The route answers at once (202, ``status: "starting"``) and the web
|
||||
process's dispatcher (web_interface/on_demand_dispatch.py) sends the
|
||||
request until the display acknowledges it; the status routes report the
|
||||
outcome. The dispatcher runs on real time with short waits here.
|
||||
"""
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def quick(self, monkeypatch, service):
|
||||
from web_interface import on_demand_dispatch
|
||||
monkeypatch.setattr(on_demand_dispatch, "RETRY_INTERVAL", 0.01)
|
||||
monkeypatch.setattr(on_demand_dispatch, "START_WAIT_SECONDS", 2.0)
|
||||
monkeypatch.setattr(f"{DISPLAY}.ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS", 1.0)
|
||||
service["cache"].get.return_value = None # the display has published nothing
|
||||
|
||||
@staticmethod
|
||||
def _outcome(status=None):
|
||||
from web_interface import on_demand_dispatch
|
||||
d = on_demand_dispatch.current()
|
||||
assert d is not None
|
||||
assert _until(lambda: not d.pending()), "the dispatcher never finished"
|
||||
return d.status()
|
||||
|
||||
def test_a_display_still_starting_gets_the_request_once_it_listens(
|
||||
self, api_v3_client, service, clock):
|
||||
self, api_v3_client, service):
|
||||
calls = []
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=_attempts(_no_socket(), _no_socket(), "ack",
|
||||
calls=calls, clock=clock)):
|
||||
side_effect=_attempts(_no_socket(), _no_socket(), "ack", calls=calls)):
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
assert resp.status_code == 200, resp.get_json()
|
||||
data = resp.get_json()["data"]
|
||||
assert data["transport"] == "socket" and "socket_error" not in data
|
||||
assert data["service"]["started"] is False
|
||||
assert resp.status_code == 202, resp.get_json()
|
||||
body = resp.get_json()
|
||||
assert body["status"] == "starting"
|
||||
data = body["data"]
|
||||
assert data["pending"] is True and data["socket_error"] == "no_socket"
|
||||
assert data["service"]["started"] is False
|
||||
outcome = self._outcome()
|
||||
assert outcome["status"] == "delivered"
|
||||
assert outcome["request_id"] == data["request_id"]
|
||||
assert len(calls) == 3
|
||||
assert not [c for c in service["calls"] if c[0] == "systemctl"]
|
||||
assert _mailbox_writes(service["cache"]) == []
|
||||
|
||||
def test_a_stopped_display_is_started_and_then_sent_the_request(
|
||||
self, api_v3_client, service, clock):
|
||||
def test_a_stopped_display_is_started_and_answered_at_once(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
calls = []
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=_attempts(*[_no_socket()] * 6, "ack",
|
||||
calls=calls, clock=clock)) as start:
|
||||
side_effect=_attempts(*[_no_socket()] * 6, "ack")) as start:
|
||||
started = time.monotonic()
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather",
|
||||
"duration": 30})
|
||||
assert resp.status_code == 200, resp.get_json()
|
||||
data = resp.get_json()["data"]
|
||||
assert data["transport"] == "socket"
|
||||
assert data["service"]["started"] is True
|
||||
answered = time.monotonic() - started
|
||||
assert resp.status_code == 202, resp.get_json()
|
||||
data = resp.get_json()["data"]
|
||||
assert data["service"]["started"] is True
|
||||
from web_interface import on_demand_dispatch
|
||||
assert data["wait_seconds"] == on_demand_dispatch.START_WAIT_SECONDS
|
||||
outcome = self._outcome()
|
||||
assert answered < 1.0, "the route waited for the display"
|
||||
assert service["calls"] == [("systemctl", "start")]
|
||||
assert len(calls) == 7
|
||||
assert outcome["status"] == "delivered"
|
||||
# The same request, the same id, every time.
|
||||
assert {call.args[0] for call in start.call_args_list} == {data["request_id"]}
|
||||
assert _mailbox_writes(service["cache"]) == []
|
||||
|
||||
def test_a_stopped_display_without_start_service_is_an_error(
|
||||
self, api_v3_client, service, clock):
|
||||
def test_while_pending_the_status_routes_say_starting(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()) as start:
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather",
|
||||
"start_service": False})
|
||||
assert resp.status_code == 400
|
||||
body = resp.get_json()
|
||||
assert "not running" in body["message"]
|
||||
assert body["data"]["socket_error"] == "no_socket"
|
||||
assert start.call_count == 1 and clock.sleeps == 0
|
||||
assert service["calls"] == []
|
||||
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
|
||||
rid = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) \
|
||||
.get_json()["data"]["request_id"]
|
||||
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
|
||||
assert status["source"] == "web"
|
||||
assert status["state"]["status"] == "starting"
|
||||
assert status["state"]["request_id"] == rid
|
||||
assert status["state"]["plugin_id"] == "weather"
|
||||
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
|
||||
assert current["on_demand_pending"]["status"] == "starting"
|
||||
|
||||
def test_a_display_that_never_comes_up_is_given_up_on(self, api_v3_client, service, clock):
|
||||
from web_interface.blueprints.api_v3 import display
|
||||
def test_a_display_that_never_comes_up_is_a_start_timeout(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
calls = []
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=_attempts(_no_socket(), calls=calls, clock=clock)):
|
||||
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
assert resp.status_code == 503
|
||||
body = resp.get_json()
|
||||
assert body["data"]["socket_error"] == "no_socket"
|
||||
assert body["data"]["service"]["started"] is True
|
||||
assert "did not answer within 45 seconds" in body["message"]
|
||||
waited = calls[-1] - calls[1]
|
||||
assert display.ON_DEMAND_SOCKET_WAIT_SECONDS - 1 <= waited \
|
||||
<= display.ON_DEMAND_SOCKET_WAIT_SECONDS
|
||||
assert _mailbox_writes(service["cache"]) == []
|
||||
assert resp.status_code == 202
|
||||
outcome = self._outcome()
|
||||
assert outcome["status"] == "error" and outcome["error"] == "start-timeout"
|
||||
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
|
||||
assert status["state"]["status"] == "error"
|
||||
assert status["state"]["error"] == "start-timeout"
|
||||
current = api_v3_client.get("/api/v3/display/current-status").get_json()["data"]
|
||||
assert current["on_demand_pending"]["error"] == "start-timeout"
|
||||
|
||||
def test_a_later_display_state_replaces_the_failure(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
|
||||
api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
failed_at = self._outcome()["last_updated"]
|
||||
later = {"active": True, "status": "active", "plugin_id": "clock",
|
||||
"last_updated": failed_at + 5}
|
||||
service["cache"].get.return_value = later
|
||||
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
|
||||
assert status["state"]["plugin_id"] == "clock" and status["source"] == "cache"
|
||||
|
||||
def test_a_running_service_without_a_socket_is_waited_for_less(
|
||||
self, api_v3_client, service, clock):
|
||||
from web_interface.blueprints.api_v3 import display
|
||||
self, api_v3_client, service):
|
||||
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
assert resp.status_code == 202
|
||||
assert resp.get_json()["data"]["wait_seconds"] == 1.0
|
||||
assert self._outcome()["error"] == "start-timeout"
|
||||
assert not [c for c in service["calls"] if c[0] == "systemctl"]
|
||||
|
||||
def test_a_different_failure_while_waiting_ends_it(self, api_v3_client, service):
|
||||
calls = []
|
||||
busy = control_client.ControlError("busy", "x", sent=True)
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=_attempts(_no_socket(), _no_socket(), busy, calls=calls)):
|
||||
assert api_v3_client.post(START_URL, json={"plugin_id": "weather"}).status_code == 202
|
||||
outcome = self._outcome()
|
||||
assert outcome["status"] == "error" and outcome["error"] == "busy"
|
||||
assert len(calls) == 3
|
||||
|
||||
def test_a_stop_while_pending_cancels_it(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
calls = []
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=_attempts(_no_socket(), calls=calls, clock=clock)):
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
assert resp.status_code == 503
|
||||
assert f"within {int(display.ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS)} seconds" \
|
||||
in resp.get_json()["message"]
|
||||
assert calls[-1] - calls[0] <= display.ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS
|
||||
assert not [c for c in service["calls"] if c[0] == "systemctl"]
|
||||
side_effect=_attempts(_no_socket(), calls=calls)), \
|
||||
patch(f"{CLIENT}.on_demand_stop", side_effect=_no_socket()):
|
||||
rid = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) \
|
||||
.get_json()["data"]["request_id"]
|
||||
resp = api_v3_client.post(STOP_URL, json={})
|
||||
assert resp.status_code == 200, resp.get_json()
|
||||
data = resp.get_json()["data"]
|
||||
assert data["cancelled_request_id"] == rid
|
||||
outcome = self._outcome()
|
||||
n = len(calls)
|
||||
time.sleep(0.05)
|
||||
assert len(calls) == n, "the cancelled start was still being sent"
|
||||
assert outcome["status"] == "idle" and outcome["last_event"] == "requested-stop"
|
||||
status = api_v3_client.get("/api/v3/display/on-demand/status").get_json()["data"]
|
||||
assert status["state"]["status"] == "idle" and status["source"] != "web"
|
||||
|
||||
def test_a_new_start_supersedes_the_pending_one(self, api_v3_client, service):
|
||||
service["state"]["active"] = False
|
||||
with patch(f"{CLIENT}.on_demand_start", side_effect=_no_socket()):
|
||||
old = api_v3_client.post(START_URL, json={"plugin_id": "weather"}) \
|
||||
.get_json()["data"]["request_id"]
|
||||
sent = []
|
||||
service["state"]["active"] = True
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=lambda rid, *a: sent.append(rid) or {"accepted": True}):
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "clock"})
|
||||
assert resp.status_code == 200
|
||||
outcome = self._outcome()
|
||||
assert sent == [resp.get_json()["data"]["request_id"]]
|
||||
assert old not in sent
|
||||
assert outcome["status"] == "idle" and outcome["last_event"] == "superseded"
|
||||
|
||||
def test_a_service_that_will_not_start_is_an_error(self, api_v3_client, service, clock):
|
||||
service["state"]["active"] = False
|
||||
@@ -233,17 +320,6 @@ class TestNoDisplayListening:
|
||||
assert "Failed to start display service" in resp.get_json()["message"]
|
||||
assert start.call_count == 1
|
||||
|
||||
def test_a_different_failure_while_waiting_ends_the_wait(self, api_v3_client, service,
|
||||
clock):
|
||||
calls = []
|
||||
busy = control_client.ControlError("busy", "x", sent=True)
|
||||
with patch(f"{CLIENT}.on_demand_start",
|
||||
side_effect=_attempts(_no_socket(), busy, calls=calls, clock=clock)):
|
||||
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
|
||||
assert resp.status_code == 503
|
||||
assert resp.get_json()["data"]["socket_error"] == "busy"
|
||||
assert len(calls) == 2
|
||||
|
||||
def test_stop_with_no_display_running_is_an_error(self, api_v3_client, service, clock):
|
||||
service["state"]["active"] = False
|
||||
with patch(f"{CLIENT}.on_demand_stop", side_effect=_no_socket()) as stop:
|
||||
|
||||
@@ -327,3 +327,27 @@ class TestCleartextIsCalledOut:
|
||||
def test_tls_on_does_not_warn(self, bridge_module):
|
||||
assert bridge_module.warn_if_cleartext(
|
||||
{"mqtt_tls": True, "mqtt_password": "hunter2"}) is False
|
||||
|
||||
|
||||
class TestTheApiClientReadsStartingAsTaken:
|
||||
"""A cold start answers 202 with ``status: "starting"``: the request is
|
||||
taken and the web process delivers it once the display listens. The
|
||||
bridge must report that as success, not as a failure."""
|
||||
|
||||
def _client(self, bridge_module, status_code, body):
|
||||
response = MagicMock(status_code=status_code)
|
||||
response.json.return_value = body
|
||||
session = MagicMock()
|
||||
session.request.return_value = response
|
||||
return bridge_module.LEDMatrixClient("http://pi:5000", session=session)
|
||||
|
||||
def test_202_starting_is_returned_not_raised(self, bridge_module):
|
||||
api = self._client(bridge_module, 202, {
|
||||
"status": "starting", "message": "starting",
|
||||
"data": {"request_id": "r1", "pending": True}})
|
||||
assert api.start_on_demand(mode="clock") == {"request_id": "r1", "pending": True}
|
||||
|
||||
def test_an_error_is_still_raised(self, bridge_module):
|
||||
api = self._client(bridge_module, 503, {"status": "error", "message": "no display"})
|
||||
with pytest.raises(RuntimeError, match="no display"):
|
||||
api.start_on_demand(mode="clock")
|
||||
|
||||
@@ -0,0 +1,249 @@
|
||||
"""The web process's on-demand dispatcher (web_interface/on_demand_dispatch.py).
|
||||
|
||||
A start that finds no display listening -- the service was just started, or
|
||||
is still loading its plugins -- is answered at once (202), and the
|
||||
dispatcher's one worker thread sends it again until the display
|
||||
acknowledges it or the wait runs out. These tests drive the worker with a
|
||||
fake ``send`` and short waits:
|
||||
|
||||
* acknowledged: delivered once, and reported as such;
|
||||
* nothing listening for the whole wait: ``start-timeout``;
|
||||
* any other failure: reported at once, not retried;
|
||||
* a newer start supersedes the pending one; a stop cancels it;
|
||||
* the outcome is reported for a while, then forgotten.
|
||||
"""
|
||||
|
||||
import threading
|
||||
import time
|
||||
|
||||
import pytest
|
||||
|
||||
from src.ipc import client as control_client
|
||||
from web_interface import on_demand_dispatch
|
||||
from web_interface.on_demand_dispatch import OnDemandDispatcher
|
||||
|
||||
|
||||
def _not_listening():
|
||||
return control_client.ControlError("no_socket", "x", sent=False)
|
||||
|
||||
|
||||
def _payload(rid, plugin_id="weather"):
|
||||
return {"request_id": rid, "action": "start", "plugin_id": plugin_id,
|
||||
"mode": plugin_id, "duration": 30, "pinned": False}
|
||||
|
||||
|
||||
class FakeSend:
|
||||
"""Answers with ``outcomes`` in turn (an exception is raised, anything
|
||||
else acks); the last one repeats. Records the request ids it was sent."""
|
||||
|
||||
def __init__(self, *outcomes):
|
||||
self.outcomes = list(outcomes) or ["ack"]
|
||||
self.sent = []
|
||||
self.lock = threading.Lock()
|
||||
|
||||
def __call__(self, payload):
|
||||
with self.lock:
|
||||
self.sent.append(payload["request_id"])
|
||||
outcome = self.outcomes.pop(0) if len(self.outcomes) > 1 else self.outcomes[0]
|
||||
if isinstance(outcome, BaseException):
|
||||
raise outcome
|
||||
if callable(outcome):
|
||||
return outcome(payload)
|
||||
return {"accepted": True}
|
||||
|
||||
|
||||
def _until(predicate, timeout=5.0):
|
||||
end = time.monotonic() + timeout
|
||||
while time.monotonic() < end:
|
||||
if predicate():
|
||||
return True
|
||||
time.sleep(0.005)
|
||||
return False
|
||||
|
||||
|
||||
def _settled(d):
|
||||
return _until(lambda: not d.pending() and d._thread is None)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def make():
|
||||
made = []
|
||||
|
||||
def build(send, wait_seconds=2.0, retry_interval=0.01):
|
||||
d = OnDemandDispatcher(send, wait_seconds=wait_seconds, retry_interval=retry_interval)
|
||||
made.append(d)
|
||||
return d
|
||||
yield build
|
||||
for d in made:
|
||||
d.cancel("teardown")
|
||||
|
||||
|
||||
class TestDelivery:
|
||||
def test_an_acknowledged_start_is_delivered_once(self, make):
|
||||
send = FakeSend("ack")
|
||||
d = make(send)
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
assert send.sent == ["r1"]
|
||||
status = d.status()
|
||||
assert status["status"] == "delivered" and status["request_id"] == "r1"
|
||||
assert status["source"] == "web"
|
||||
|
||||
def test_it_is_sent_again_until_the_display_listens(self, make):
|
||||
send = FakeSend(_not_listening(), _not_listening(), "ack")
|
||||
d = make(send)
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
assert send.sent == ["r1", "r1", "r1"]
|
||||
assert d.status()["status"] == "delivered"
|
||||
|
||||
def test_while_it_waits_it_reports_starting(self, make):
|
||||
gate = threading.Event()
|
||||
|
||||
def blocked(payload):
|
||||
gate.wait(5)
|
||||
raise _not_listening()
|
||||
|
||||
send = FakeSend(blocked, "ack")
|
||||
d = make(send)
|
||||
d.submit(_payload("r1"))
|
||||
status = d.status()
|
||||
assert status["status"] == "starting" and status["active"] is False
|
||||
assert (status["plugin_id"], status["mode"], status["duration"]) == ("weather", "weather", 30)
|
||||
gate.set()
|
||||
assert _settled(d)
|
||||
assert d.status()["status"] == "delivered"
|
||||
|
||||
|
||||
class TestGivingUp:
|
||||
def test_nothing_listening_for_the_whole_wait_is_a_start_timeout(self, make):
|
||||
send = FakeSend(_not_listening())
|
||||
d = make(send, wait_seconds=0.2)
|
||||
started = time.monotonic()
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
waited = time.monotonic() - started
|
||||
status = d.status()
|
||||
assert status["status"] == "error" and status["error"] == "start-timeout"
|
||||
assert 0.15 <= waited < 2.0
|
||||
assert len(send.sent) > 3 # it kept trying in between
|
||||
|
||||
def test_a_start_can_carry_its_own_wait(self, make):
|
||||
# The route passes a shorter wait for a service that was already
|
||||
# running; the dispatcher's default must not override it.
|
||||
d = make(FakeSend(_not_listening()), wait_seconds=30.0)
|
||||
d.submit(_payload("r1"), wait_seconds=0.1)
|
||||
assert _until(lambda: not d.pending(), timeout=3.0), "it waited the default"
|
||||
assert d.status()["error"] == "start-timeout"
|
||||
|
||||
@pytest.mark.parametrize("reason,sent", [("busy", True), ("unknown_command", True),
|
||||
("timeout", True), ("forbidden", False)])
|
||||
def test_any_other_failure_is_reported_at_once(self, make, reason, sent):
|
||||
send = FakeSend(control_client.ControlError(reason, "x", sent=sent))
|
||||
d = make(send)
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
assert send.sent == ["r1"]
|
||||
status = d.status()
|
||||
assert status["status"] == "error" and status["error"] == reason
|
||||
|
||||
def test_a_client_bug_is_internal(self, make):
|
||||
d = make(FakeSend(RuntimeError("boom")))
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
assert d.status()["error"] == "internal"
|
||||
|
||||
|
||||
class TestOneAtATime:
|
||||
def test_a_newer_start_supersedes_the_pending_one(self, make):
|
||||
send = FakeSend(_not_listening())
|
||||
d = make(send)
|
||||
d.submit(_payload("old"))
|
||||
assert _until(lambda: "old" in send.sent)
|
||||
d.submit(_payload("new", plugin_id="clock"))
|
||||
assert d.status()["request_id"] == "new"
|
||||
n = len(send.sent)
|
||||
send.outcomes = ["ack"]
|
||||
assert _settled(d)
|
||||
assert send.sent[n:] and set(send.sent[n + 1:]) <= {"new"}
|
||||
assert send.sent[-1] == "new"
|
||||
status = d.status()
|
||||
assert status["status"] == "delivered" and status["plugin_id"] == "clock"
|
||||
|
||||
def test_an_answer_for_a_superseded_start_is_not_reported(self, make):
|
||||
# The old start's send is in flight when the new one arrives; its
|
||||
# ack must not mark the new one delivered.
|
||||
in_flight, release = threading.Event(), threading.Event()
|
||||
|
||||
def slow(payload):
|
||||
in_flight.set()
|
||||
release.wait(5)
|
||||
return {"accepted": True}
|
||||
|
||||
send = FakeSend(slow, _not_listening())
|
||||
d = make(send, wait_seconds=0.3)
|
||||
d.submit(_payload("old"))
|
||||
assert in_flight.wait(5)
|
||||
d.submit(_payload("new"))
|
||||
release.set()
|
||||
assert _settled(d)
|
||||
status = d.status()
|
||||
assert status["request_id"] == "new"
|
||||
assert status["status"] == "error" and status["error"] == "start-timeout"
|
||||
|
||||
def test_a_stop_cancels_the_pending_start(self, make):
|
||||
send = FakeSend(_not_listening())
|
||||
d = make(send)
|
||||
d.submit(_payload("r1"))
|
||||
assert _until(lambda: send.sent)
|
||||
assert d.cancel("requested-stop") == "r1"
|
||||
assert _settled(d)
|
||||
n = len(send.sent)
|
||||
time.sleep(0.05)
|
||||
assert len(send.sent) == n, "it kept sending a cancelled start"
|
||||
status = d.status()
|
||||
assert status["status"] == "idle" and status["last_event"] == "requested-stop"
|
||||
|
||||
def test_a_cancel_with_nothing_pending_does_nothing(self, make):
|
||||
d = make(FakeSend("ack"))
|
||||
assert d.cancel() is None
|
||||
assert d.status() is None
|
||||
|
||||
def test_a_cancel_after_delivery_leaves_the_outcome(self, make):
|
||||
d = make(FakeSend("ack"))
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
assert d.cancel() is None
|
||||
assert d.status()["status"] == "delivered"
|
||||
|
||||
|
||||
class TestOutcomeLifetime:
|
||||
def test_an_outcome_is_forgotten_after_a_while(self, make, monkeypatch):
|
||||
d = make(FakeSend("ack"))
|
||||
d.submit(_payload("r1"))
|
||||
assert _settled(d)
|
||||
assert d.status() is not None
|
||||
monkeypatch.setattr(on_demand_dispatch, "OUTCOME_SECONDS", 0.0)
|
||||
time.sleep(0.01)
|
||||
assert d.status() is None
|
||||
|
||||
def test_a_pending_start_never_expires(self, make, monkeypatch):
|
||||
monkeypatch.setattr(on_demand_dispatch, "OUTCOME_SECONDS", 0.0)
|
||||
gate = threading.Event()
|
||||
d = make(FakeSend(lambda p: gate.wait(5) and {"accepted": True}))
|
||||
d.submit(_payload("r1"))
|
||||
time.sleep(0.01)
|
||||
assert d.status()["status"] == "starting"
|
||||
gate.set()
|
||||
assert _settled(d)
|
||||
|
||||
|
||||
def test_the_process_has_one_dispatcher():
|
||||
on_demand_dispatch.reset_for_tests()
|
||||
try:
|
||||
assert on_demand_dispatch.current() is None
|
||||
first = on_demand_dispatch.get_dispatcher(FakeSend())
|
||||
assert on_demand_dispatch.get_dispatcher(FakeSend()) is first
|
||||
assert on_demand_dispatch.current() is first
|
||||
finally:
|
||||
on_demand_dispatch.reset_for_tests()
|
||||
@@ -9,7 +9,7 @@ from web_interface.blueprints.api_v3 import (
|
||||
_get_display_service_status, _socket_reason_code, _stop_display_service, api_v3,
|
||||
jsonify, logger, request, uuid,
|
||||
)
|
||||
from web_interface import display_preview, display_state
|
||||
from web_interface import display_preview, display_state, on_demand_dispatch
|
||||
import web_interface.blueprints.api_v3 as _pkg
|
||||
from src.ipc import client as control_client
|
||||
# Read through the module rather than bound by value: tests patch these
|
||||
@@ -33,17 +33,29 @@ def _cache_manager():
|
||||
|
||||
|
||||
|
||||
#: How long the start route waits for a display it has just started (or one
|
||||
#: systemd already reports running, which may still be loading its plugins)
|
||||
#: to serve its control socket, before it gives up. The socket comes up when
|
||||
#: the display's run loop starts, after every plugin has loaded.
|
||||
ON_DEMAND_SOCKET_WAIT_SECONDS = 45.0
|
||||
#: The same wait when the service was already running: a display that has
|
||||
#: just been restarted by someone else. Shorter, because a running display
|
||||
#: normally has its socket.
|
||||
#: How long a start is sent again to a service that systemd reports running
|
||||
#: but that has no socket yet (a display still loading its plugins, or one
|
||||
#: someone else just restarted). A cold start gets the dispatcher's own
|
||||
#: START_WAIT_SECONDS. Either way the route answers at once (202) and the
|
||||
#: web process's dispatcher does the waiting.
|
||||
ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS = 10.0
|
||||
#: Gap between two attempts while waiting for the socket.
|
||||
ON_DEMAND_SOCKET_RETRY_INTERVAL = 0.5
|
||||
|
||||
|
||||
def _dispatcher():
|
||||
"""The web process's on-demand dispatcher (web_interface/on_demand_dispatch.py)."""
|
||||
return on_demand_dispatch.get_dispatcher(_send_on_demand)
|
||||
|
||||
|
||||
def _pending_start_state():
|
||||
"""A start the dispatcher is still delivering, or one it gave up on:
|
||||
the state the status routes report instead of the display's. None when
|
||||
there is none (or it was delivered, after which the display's own
|
||||
state is the truth)."""
|
||||
dispatcher = on_demand_dispatch.current()
|
||||
status = dispatcher.status() if dispatcher is not None else None
|
||||
if status is None or status.get('status') not in ('starting', 'error'):
|
||||
return None
|
||||
return status
|
||||
|
||||
|
||||
def _send_on_demand(payload):
|
||||
@@ -61,21 +73,6 @@ def _send_on_demand(payload):
|
||||
return control_client.on_demand_stop(payload['request_id'])
|
||||
|
||||
|
||||
def _send_on_demand_when_listening(payload, wait_seconds):
|
||||
"""_send_on_demand, retried while no display is listening yet (a display
|
||||
still starting), for up to ``wait_seconds``. Any other failure, or the
|
||||
last one once the time is up, raises ``control_client.ControlError``."""
|
||||
deadline = _pkg.time.monotonic() + wait_seconds
|
||||
while True:
|
||||
try:
|
||||
return _send_on_demand(payload)
|
||||
except control_client.ControlError as e:
|
||||
if (not control_client.display_not_listening(e)
|
||||
or _pkg.time.monotonic() + ON_DEMAND_SOCKET_RETRY_INTERVAL > deadline):
|
||||
raise
|
||||
_pkg.time.sleep(ON_DEMAND_SOCKET_RETRY_INTERVAL)
|
||||
|
||||
|
||||
def _socket_error_response(request_id, action, reason, message=None, **extra):
|
||||
"""The answer when the display did not take an on-demand request: ``400``
|
||||
for arguments it refused, else ``503``."""
|
||||
@@ -205,6 +202,7 @@ def get_on_demand_status():
|
||||
"""
|
||||
state = display_state.on_demand_state(display_state.read_state())
|
||||
source = 'socket'
|
||||
pending = _pending_start_state()
|
||||
if state is None:
|
||||
source = 'cache'
|
||||
cache = _cache_manager()
|
||||
@@ -213,6 +211,10 @@ def get_on_demand_status():
|
||||
# copy it read for the full max_age -- "active" for two minutes after
|
||||
# the display had already stopped.
|
||||
state = cache.get('display_on_demand_state', max_age=120, memory_ttl=0)
|
||||
if pending is not None and _shadows(pending, state):
|
||||
# A start the web process is still delivering, or gave up on
|
||||
# (start-timeout): newer than anything the display has said.
|
||||
state, source = pending, 'web'
|
||||
if state is None:
|
||||
state = {
|
||||
'active': False,
|
||||
@@ -228,6 +230,18 @@ def get_on_demand_status():
|
||||
'source': source,
|
||||
}
|
||||
})
|
||||
def _shadows(pending, state):
|
||||
"""Whether the web process's pending start (or its failure) is newer
|
||||
than the display's on-demand ``state``. While it is still being sent it
|
||||
always is; a failure is, until the display publishes something later."""
|
||||
if pending.get('status') == 'starting' or not isinstance(state, dict):
|
||||
return True
|
||||
shown = state.get('last_updated')
|
||||
if not isinstance(shown, (int, float)) or isinstance(shown, bool):
|
||||
return True
|
||||
return shown < (pending.get('last_updated') or 0)
|
||||
|
||||
|
||||
@api_v3.route('/display/on-demand/start', methods=['POST'])
|
||||
def start_on_demand_display():
|
||||
"""Request the display controller to run a specific plugin on-demand."""
|
||||
@@ -290,6 +304,11 @@ def start_on_demand_display():
|
||||
'pinned': pinned,
|
||||
'timestamp': _pkg.time.time()
|
||||
}
|
||||
# This start supersedes one the dispatcher is still delivering,
|
||||
# whatever becomes of it: an older request must not land after it.
|
||||
dispatcher = on_demand_dispatch.current()
|
||||
if dispatcher is not None:
|
||||
dispatcher.cancel('superseded')
|
||||
try:
|
||||
_send_on_demand(request_payload)
|
||||
except Exception as e: # pylint: disable=broad-except
|
||||
@@ -325,8 +344,10 @@ def start_on_demand_display():
|
||||
# start_service means "start it if it is not running", as the UI's
|
||||
# checkbox says; _ensure_display_service_running leaves a running
|
||||
# service alone (restarting it cost seconds of blank panel for nothing).
|
||||
# Either way the display has no socket yet: wait for it, then send the
|
||||
# request again.
|
||||
# Either way the display has no socket yet. The route does not wait for
|
||||
# it -- a cold start can outlast a client's timeout (the MQTT bridge's is
|
||||
# 15 s) -- but answers 202 and leaves the sending to the dispatcher,
|
||||
# whose outcome the status routes report.
|
||||
wait = ON_DEMAND_SOCKET_WAIT_RUNNING_SECONDS
|
||||
service_result = None
|
||||
if not service_status.get('active'):
|
||||
@@ -337,24 +358,28 @@ def start_on_demand_display():
|
||||
'message': 'Failed to start display service. Please check service logs or start it manually.',
|
||||
'service_result': service_result
|
||||
}), 500
|
||||
wait = ON_DEMAND_SOCKET_WAIT_SECONDS
|
||||
wait = on_demand_dispatch.START_WAIT_SECONDS
|
||||
elif start_service:
|
||||
service_result = dict(service_status, started=False)
|
||||
|
||||
try:
|
||||
_send_on_demand_when_listening(request_payload, wait)
|
||||
except Exception as e: # pylint: disable=broad-except
|
||||
reason = _socket_failure_reason(e)
|
||||
if control_client.display_not_listening(e):
|
||||
message = (f'The display service is running but its control socket did not '
|
||||
f'answer within {int(wait)} seconds ({reason}). It may still be '
|
||||
f'starting; try again shortly, or check its logs.')
|
||||
return _socket_error_response(request_id, 'start', reason, message,
|
||||
service=service_result)
|
||||
return _socket_error_response(request_id, 'start', reason, service=service_result)
|
||||
|
||||
return _on_demand_started(request_id, resolved_plugin, resolved_mode,
|
||||
duration, pinned, service_result)
|
||||
_dispatcher().submit(request_payload, wait_seconds=wait)
|
||||
return jsonify({
|
||||
'status': 'starting',
|
||||
'message': ('The display service is starting; the request is sent to it as soon '
|
||||
'as it is listening. Check the on-demand status for the outcome.'),
|
||||
'data': {
|
||||
'request_id': request_id,
|
||||
'plugin_id': resolved_plugin,
|
||||
'mode': resolved_mode,
|
||||
'duration': duration,
|
||||
'pinned': pinned,
|
||||
'service': service_result,
|
||||
'transport': 'socket',
|
||||
'socket_error': reason,
|
||||
'pending': True,
|
||||
'wait_seconds': wait,
|
||||
},
|
||||
}), 202
|
||||
|
||||
|
||||
def _on_demand_started(request_id, plugin_id, mode, duration, pinned, service_result):
|
||||
@@ -386,11 +411,15 @@ def stop_on_demand_display():
|
||||
'timestamp': _pkg.time.time()
|
||||
}
|
||||
socket_error = None
|
||||
# A start the web process is still delivering is dropped first: the
|
||||
# stop is newer, whatever happens to it below.
|
||||
dispatcher = on_demand_dispatch.current()
|
||||
cancelled = dispatcher.cancel('requested-stop') if dispatcher is not None else None
|
||||
try:
|
||||
_send_on_demand(request_payload)
|
||||
except Exception as e: # pylint: disable=broad-except
|
||||
socket_error = _socket_failure_reason(e)
|
||||
if not stop_service:
|
||||
if not stop_service and not (cancelled and control_client.display_not_listening(e)):
|
||||
if control_client.display_not_listening(e):
|
||||
service_status = _get_display_service_status()
|
||||
message = ('Display service is not running, so the stop could not be '
|
||||
@@ -414,6 +443,9 @@ def stop_on_demand_display():
|
||||
'service': service_result,
|
||||
'transport': 'socket',
|
||||
}
|
||||
if cancelled:
|
||||
# The start it ended never reached the display.
|
||||
response_data['cancelled_request_id'] = cancelled
|
||||
if socket_error:
|
||||
response_data['socket_error'] = socket_error
|
||||
return jsonify({'status': 'success', 'data': response_data})
|
||||
@@ -448,4 +480,10 @@ def get_current_display_status():
|
||||
'plugin_id': None,
|
||||
'last_updated': None,
|
||||
}
|
||||
return jsonify({'status': 'success', 'data': dict(state, source=source)})
|
||||
data = dict(state, source=source)
|
||||
pending = _pending_start_state()
|
||||
if pending is not None:
|
||||
# An on-demand start the web process is still delivering (or gave
|
||||
# up on): what the panel is about to show, or why it will not.
|
||||
data['on_demand_pending'] = pending
|
||||
return jsonify({'status': 'success', 'data': data})
|
||||
|
||||
@@ -0,0 +1,221 @@
|
||||
"""Deliver an on-demand start to a display that is not listening yet.
|
||||
|
||||
``POST /api/v3/display/on-demand/start`` can find no display on the control
|
||||
socket: the service is stopped (the route starts it) or still loading its
|
||||
plugins. The socket comes up only when the display's run loop starts, which
|
||||
can take longer than a client waits -- the MQTT bridge gives up after 15 s.
|
||||
So the route answers at once (``202``, ``status: "starting"``) and hands
|
||||
the request to the one :class:`OnDemandDispatcher` of the web process, whose
|
||||
worker thread sends it again until the display acknowledges it or
|
||||
:data:`START_WAIT_SECONDS` pass.
|
||||
|
||||
* One request at a time: a newer start replaces the pending one, and a
|
||||
stop cancels it (:meth:`OnDemandDispatcher.cancel`).
|
||||
* Its outcome is :meth:`OnDemandDispatcher.status`, which
|
||||
``GET /display/on-demand/status`` and ``/display/current-status`` report:
|
||||
``starting`` while it waits, ``delivered`` once acknowledged (the
|
||||
display's own state takes over from there), or ``error`` with
|
||||
``start-timeout`` or the socket's reason.
|
||||
|
||||
Nothing is written to disk: the file mailbox that once carried such a
|
||||
request is gone (docs/IPC_CONTROL_SOCKET.md, stage 5).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import threading
|
||||
import time
|
||||
from typing import Any, Callable, Dict, Optional
|
||||
|
||||
from src.ipc import client as control_client
|
||||
from src.logging_config import get_logger
|
||||
|
||||
logger = get_logger(__name__)
|
||||
|
||||
#: How long the worker keeps sending a start before it gives up
|
||||
#: (``start-timeout``). The socket comes up when the display's run loop
|
||||
#: starts, after every plugin has loaded.
|
||||
START_WAIT_SECONDS = 45.0
|
||||
|
||||
#: Gap between two sends while nothing is listening.
|
||||
RETRY_INTERVAL = 0.5
|
||||
|
||||
#: How long a finished outcome (delivered, error, cancelled) is still
|
||||
#: reported, so a client polling every few seconds sees it.
|
||||
OUTCOME_SECONDS = 120.0
|
||||
|
||||
#: ``send(payload)`` hands the request to the display (the route's
|
||||
#: ``_send_on_demand``) and raises ``ControlError`` when it does not take it.
|
||||
Sender = Callable[[Dict[str, Any]], Any]
|
||||
|
||||
|
||||
class OnDemandDispatcher:
|
||||
"""One pending on-demand start, and the worker thread that delivers it."""
|
||||
|
||||
def __init__(self, send: Sender, *,
|
||||
wait_seconds: Optional[float] = None,
|
||||
retry_interval: Optional[float] = None,
|
||||
clock: Callable[[], float] = time.monotonic,
|
||||
wall_clock: Callable[[], float] = time.time):
|
||||
self._send = send
|
||||
self.wait_seconds = START_WAIT_SECONDS if wait_seconds is None else wait_seconds
|
||||
self.retry_interval = RETRY_INTERVAL if retry_interval is None else retry_interval
|
||||
self._clock = clock
|
||||
self._wall = wall_clock
|
||||
self._lock = threading.Lock()
|
||||
self._wake = threading.Event()
|
||||
# Bumped by every submit and cancel: a send that started under an
|
||||
# older generation does not report its result as the current one.
|
||||
self._generation = 0
|
||||
self._pending: Optional[Dict[str, Any]] = None
|
||||
self._deadline = 0.0
|
||||
self._status: Optional[Dict[str, Any]] = None
|
||||
self._finished_at: Optional[float] = None
|
||||
self._thread: Optional[threading.Thread] = None
|
||||
|
||||
# -- the routes' side ------------------------------------------------------
|
||||
|
||||
def submit(self, payload: Dict[str, Any], wait_seconds: Optional[float] = None) -> None:
|
||||
"""Deliver ``payload`` (an on-demand start) in the background,
|
||||
replacing any start still pending."""
|
||||
with self._lock:
|
||||
self._generation += 1
|
||||
superseded = self._pending
|
||||
self._pending = dict(payload)
|
||||
self._deadline = self._clock() + (self.wait_seconds if wait_seconds is None
|
||||
else wait_seconds)
|
||||
self._status = self._describe('starting', payload)
|
||||
self._finished_at = None
|
||||
if self._thread is None or not self._thread.is_alive():
|
||||
self._thread = threading.Thread(target=self._run, name='on-demand-dispatch',
|
||||
daemon=True)
|
||||
self._thread.start()
|
||||
if superseded is not None:
|
||||
logger.info("On-demand start %s superseded by %s before the display took it",
|
||||
superseded.get('request_id'), payload.get('request_id'))
|
||||
self._wake.set()
|
||||
|
||||
def cancel(self, reason: str = 'cancelled') -> Optional[str]:
|
||||
"""Drop the pending start (a stop arrived). Returns its request id,
|
||||
or None when nothing was pending."""
|
||||
with self._lock:
|
||||
pending = self._pending
|
||||
if pending is None:
|
||||
return None
|
||||
self._generation += 1
|
||||
self._pending = None
|
||||
self._finish(self._describe('idle', pending, last_event=reason))
|
||||
self._wake.set()
|
||||
logger.info("On-demand start %s cancelled before the display took it (%s)",
|
||||
pending.get('request_id'), reason)
|
||||
return pending.get('request_id')
|
||||
|
||||
def status(self) -> Optional[Dict[str, Any]]:
|
||||
"""The pending start's state, or its outcome for OUTCOME_SECONDS
|
||||
after it finished; None otherwise. In the shape of the display's
|
||||
on-demand state (``active``, ``status``, ``error``, ...), plus
|
||||
``source: "web"``."""
|
||||
with self._lock:
|
||||
if self._status is None:
|
||||
return None
|
||||
if (self._finished_at is not None
|
||||
and self._clock() - self._finished_at > OUTCOME_SECONDS):
|
||||
return None
|
||||
return dict(self._status)
|
||||
|
||||
def pending(self) -> bool:
|
||||
with self._lock:
|
||||
return self._pending is not None
|
||||
|
||||
# -- the worker ------------------------------------------------------------
|
||||
|
||||
def _describe(self, status: str, payload: Dict[str, Any], error: Optional[str] = None,
|
||||
last_event: Optional[str] = None) -> Dict[str, Any]:
|
||||
return {
|
||||
'active': False,
|
||||
'status': status,
|
||||
'error': error,
|
||||
'last_event': last_event,
|
||||
'request_id': payload.get('request_id'),
|
||||
'plugin_id': payload.get('plugin_id'),
|
||||
'mode': payload.get('mode'),
|
||||
'duration': payload.get('duration'),
|
||||
'pinned': bool(payload.get('pinned', False)),
|
||||
'last_updated': self._wall(),
|
||||
'source': 'web',
|
||||
}
|
||||
|
||||
def _finish(self, status: Dict[str, Any]) -> None:
|
||||
"""Record an outcome. Caller holds _lock."""
|
||||
self._status = status
|
||||
self._finished_at = self._clock()
|
||||
|
||||
def _run(self) -> None:
|
||||
while True:
|
||||
with self._lock:
|
||||
payload, generation = self._pending, self._generation
|
||||
deadline = self._deadline
|
||||
if payload is None:
|
||||
self._thread = None
|
||||
return
|
||||
outcome, error = self._attempt(payload)
|
||||
with self._lock:
|
||||
if generation != self._generation:
|
||||
continue # superseded or cancelled meanwhile
|
||||
if outcome == 'retry' and self._clock() + self.retry_interval > deadline:
|
||||
outcome, error = 'error', 'start-timeout'
|
||||
if outcome == 'delivered':
|
||||
self._pending = None
|
||||
self._finish(self._describe('delivered', payload,
|
||||
last_event='delivered'))
|
||||
elif outcome == 'error':
|
||||
self._pending = None
|
||||
self._finish(self._describe('error', payload, error=error))
|
||||
else:
|
||||
self._wake.clear()
|
||||
if outcome == 'delivered':
|
||||
logger.info("On-demand start %s delivered once the display was listening",
|
||||
payload.get('request_id'))
|
||||
elif outcome == 'error':
|
||||
logger.warning("On-demand start %s not delivered: %s",
|
||||
payload.get('request_id'), error)
|
||||
else:
|
||||
# A submit or cancel wakes the wait at once.
|
||||
self._wake.wait(self.retry_interval)
|
||||
|
||||
def _attempt(self, payload: Dict[str, Any]):
|
||||
try:
|
||||
self._send(payload)
|
||||
except control_client.ControlError as e:
|
||||
if control_client.display_not_listening(e):
|
||||
return 'retry', None
|
||||
return 'error', str(e.reason)
|
||||
except Exception: # pylint: disable=broad-except
|
||||
logger.exception("On-demand start %s: the control socket client failed",
|
||||
payload.get('request_id'))
|
||||
return 'error', 'internal'
|
||||
return 'delivered', None
|
||||
|
||||
|
||||
_dispatcher: Optional[OnDemandDispatcher] = None
|
||||
_dispatcher_lock = threading.Lock()
|
||||
|
||||
|
||||
def get_dispatcher(send: Sender) -> OnDemandDispatcher:
|
||||
"""The web process's dispatcher, created on first use."""
|
||||
global _dispatcher
|
||||
with _dispatcher_lock:
|
||||
if _dispatcher is None:
|
||||
_dispatcher = OnDemandDispatcher(send)
|
||||
return _dispatcher
|
||||
|
||||
|
||||
def current() -> Optional[OnDemandDispatcher]:
|
||||
"""The dispatcher if one was created, without creating one."""
|
||||
return _dispatcher
|
||||
|
||||
|
||||
def reset_for_tests() -> None:
|
||||
global _dispatcher
|
||||
with _dispatcher_lock:
|
||||
_dispatcher = None
|
||||
@@ -356,9 +356,12 @@ window.previewPluginNow = function(pluginId) {
|
||||
})
|
||||
.then(r => r.json())
|
||||
.then(data => {
|
||||
// 'starting' (202): the display service is starting and the request
|
||||
// follows once it listens.
|
||||
const starting = data.status === 'starting';
|
||||
showNotification(data.message || ('Previewing ' + pluginId + ' for 60 seconds'),
|
||||
data.status || 'success');
|
||||
if (data.status === 'success') window.toggleFloatingPreview(true);
|
||||
starting ? 'info' : (data.status || 'success'));
|
||||
if (data.status === 'success' || starting) window.toggleFloatingPreview(true);
|
||||
})
|
||||
.catch(err => {
|
||||
showNotification('Preview failed: ' + err.message, 'error');
|
||||
|
||||
@@ -1739,6 +1739,13 @@ function submitOnDemandRequest(event) {
|
||||
showNotification(`Requested on-demand mode for ${pluginName}`, 'success');
|
||||
closeOnDemandModal();
|
||||
setTimeout(() => loadOnDemandStatus(true), 700);
|
||||
} else if (result.status === 'starting') {
|
||||
// 202: the display service is starting; the web process sends
|
||||
// the request once it listens. The status card shows the outcome.
|
||||
const pluginName = resolvePluginDisplayName(currentOnDemandPluginId);
|
||||
showNotification(`Starting the display for ${pluginName}…`, 'info');
|
||||
closeOnDemandModal();
|
||||
setTimeout(() => loadOnDemandStatus(true), 700);
|
||||
} else {
|
||||
console.error('[submitOnDemandRequest] Request failed:', result);
|
||||
showNotification(result.message || 'Failed to start on-demand mode', 'error');
|
||||
|
||||
Reference in New Issue
Block a user