On ledpi (three cold starts) the display acknowledged the start as its
socket opened, then took ~5 s to act on it while Vegas built its first
strip; the status routes meanwhile showed the display's own idle state, so
a UI polling every 700 ms flashed idle.
The display's on-demand state now names the request it answers
(request_id). A delivered start keeps reading as status "starting" with
delivered: true, in /display/on-demand/status and as on_demand_pending in
/display/current-status, until the display publishes state for that
request id (a display without the field: any state newer than the
delivery), for at most DELIVERED_SHOWN_SECONDS (30 s). The display's
startup state, which can be published after the acknowledgement, names no
request and does not end it.
Tests: stays starting against the startup idle state (no id, an older id);
the matching active state and the matching error take over; an older
display's newer state takes over; the 30 s cap; the display's state names
its request. Mutation check: 11 mutants, 11 killed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI caught a race in test_a_new_start_supersedes_the_pending_one: the
dispatcher's worker could read the old start, the route cancel it and send
the new one, and the worker's send of the old one land after it. cancel()
now waits out a send already in flight (the worker holds a send lock while
it reads the pending start and sends it), so whatever the caller sends next
lands after it. Pinned by test_a_cancel_waits_for_a_send_in_flight, which
fails without the wait.
The Linux-only TestRealSocket test for a display that went away now
expects the 202 and the start-timeout that follows, as the route answers
since the background dispatcher.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The start route held a request open for up to 45 s while a cold-started
display loaded its plugins; the MQTT bridge (15 s timeout) and browsers
reported a failure for a request that was then delivered.
Now, when no display is listening, the route starts the service if asked
and answers 202 with status "starting" at once. A single worker in the web
process (web_interface/on_demand_dispatch.py) sends the request until the
display acknowledges it or the wait runs out (45 s cold start, 10 s for a
running service without a socket yet). A newer start supersedes the
pending one; a stop cancels it (and succeeds, with cancelled_request_id,
even with no display listening). The outcome is reported by
/display/on-demand/status (source "web": starting, or error with
start-timeout or the socket's reason, until the display publishes
something newer) and by /display/current-status as on_demand_pending.
Callers: the web UI's on-demand modal and "Preview on display" treat
"starting" as taken (an info toast); the MQTT bridge already treats any
non-error 2xx as success (now pinned by a test).
Tests: the dispatcher (ack, retry then ack, start-timeout, other failures,
superseded, an in-flight ack for a superseded start, stop while pending,
a per-start wait, outcome lifetime); the routes (202, status routes while
pending and after a timeout, a later display state replacing the failure,
stop while pending, a new start superseding); a JS suite for app.js.
Mutation check: 20 mutants on the worker, the routes and app.js, 20 killed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>