Compare commits

..
Author SHA1 Message Date
ChuckandClaude Opus 5.5 e39dcb18b4 feat(scroll): report a panel that cannot reach its refresh cap, and suggest one it can hold
Scroll speeds are solved against display.hardware.limit_refresh_rate_hz,
which is only a ceiling. A panel that cannot reach it still moves whole
pixels per frame, but every scroll runs slow by the shortfall and the
"smooth" ladder is the cap's, not the panel's. A user rig (Pi 4, 2x128x64,
adafruit-hat-pwm, pwm_bits 9, gpio_slowdown 5) measured 107.6-113.1 Hz under
a 120 Hz cap: 60 px/s ran at 55, and nothing said why.

- scroll_config: refresh_shortfall() (more than 3% under the planned rate),
  holdable_cap() (a multiple of 10, 5% under the measurement, since the
  measurement is the fast end of an uncapped panel's drift), and
  describe_refresh_shortfall().
- FrameTimingRecorder.plan_refresh(): once the measured period has held for
  three trusted windows, a shortfall is logged once as a warning naming the
  cap to use. DisplayManager calls it only for a real panel, not the
  emulator or the fallback canvas. The stats file records
  planned_refresh_hz (additive).
- GET /api/v3/config/refresh-rate, plus a hint under the Display tab's
  Limit Refresh Rate field with a button that fills in the suggested cap.
- _panel_refresh_hz (behind the Vegas slider's advice) ignores a measurement
  written under a different cap, so a changed cap stops being advised from
  the old rate before the display restarts.

Verified on ledpi with a temporary 200 Hz cap: the warning logged about a
minute after the restart ("about 132 Hz ... Set Limit Refresh Rate to
120 Hz"), the endpoint returned the same shortfall, and the Display tab
showed the hint; its button filled in 120. ledpi was restored afterwards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 17:11:09 -04:00
54 changed files with 1232 additions and 4781 deletions
+17 -111
View File
@@ -19,94 +19,24 @@ accepts both, but the store flags the old spelling as deprecated
## Unreleased
### The control socket carries every web command; the mailboxes are a fallback
### Scroll speed: a panel slower than its refresh cap is reported
Stage 4 of the web → display control socket (`docs/IPC_CONTROL_SOCKET.md`).
- **Mailbox only when the socket cannot carry it.** The on-demand routes
(`POST /api/v3/display/on-demand/start` and `/stop`) write the
`display_on_demand_request` mailbox only when the display never had the
request: no socket (a stopped display, one older than the socket), a
refused or timed-out connect, or a display too old to know the command.
A display that had it and refused or did not answer (a full queue, bad
arguments, silence after the send) is answered `503` (`400` for bad
arguments) with `socket_error`, and no mailbox copy is written: the
display may have applied it, or would refuse the copy too. A stop with
`stop_service` still stops the service. `src.ipc.client.should_fall_back()`
holds the rule; `ControlError.sent` says whether the display had the
request.
- **`errors.clear`.** `POST /api/v3/errors/clear` goes over the socket: the
display clears its error records and republishes its error snapshot
before it answers, so the response says `applied: true` with the
display's own `cleared_count`. The `plugin_error_clear_request` mailbox
is written only on the same fallback rule (a display from before this
release answers `unknown_command`, and gets the mailbox). A display that
had it and failed answers `503`. The error snapshot gains
`applied_clear_cutoff`, so an older mailbox request is not shown as
pending once a wider clear has been applied.
- **The display looks at the mailboxes less, and more cheaply.** While the
control socket is up, the on-demand mailbox is looked at once a second
instead of every 0.25 s (`MAILBOX_POLL_INTERVAL_WITH_SOCKET`), and both
mailboxes are read only when their file changed since the last look:
otherwise a look is one `stat()` (`CacheManager.file_signature`,
`MailboxWatch`). A socket command no longer reads or deletes the mailbox
file. A duplicate already processed is taken out of the mailbox, rather
than re-read for an hour. Without a socket (Windows,
`LEDMATRIX_CONTROL_SOCKET=off`) the mailbox is read every 0.25 s as before.
- **Kept for one release.** The display still reads both mailboxes, so an
older web interface (or a web user not yet in the socket's group) keeps
working during an upgrade, and still writes `display_current_state`,
`display_on_demand_state` and `plugin_runtime_snapshot` for the readers'
fallback. A request that comes through the on-demand mailbox while the
socket is up is logged once per writer: plugins that write
`display_on_demand_request` themselves (birdnet-go, mqtt-notifications,
on-air, pomodoro-timer) now get the screen within a second rather than a
quarter second, and need an in-process way in before the mailbox goes.
### Display loop stage 3: a ScreenRunner, and the Arbiter decides every screen
Internal; no behaviour change. Stage 3 of `docs/RUN_LOOP_REDESIGN.md`.
- Each screen runs in `ScreenRunner` (`src/screen_runner.py`): the first
frame, the 125 Hz or 1 Hz frame loop, the make-up dwell and the
dynamic-duration exit, moved out of `DisplayController.run()` with their
pacing unchanged. It paces with an injected clock and returns an
`Outcome` whose `ExitReason` is `DURATION`, `CYCLE_COMPLETE`, `EMPTY`,
`ERROR`, `DISPLAY_FALSE`, `RELOAD` or `PREEMPTED`. `PREEMPTED` replaces
the five "did the mode change under this screen?" re-checks.
- `Arbiter.decide()` now answers for on-demand, live priority and the
rotation too (Sources `ON_DEMAND`, `LIVE`, `ROTATION`); `LEGACY` means
only Vegas, whose iteration moves to stage 4. The on-demand session, the
rotation's position and the live resume point are snapshotted into
`ArbiterState`, whose pure transitions (`next_on_demand`, `claim_live`,
`release_live`, `after`) replace the bookkeeping in `_resolve_active_mode`,
`_apply_live_priority` and `_advance_after_screen`.
- Between frames, the runner's service points make one
`decide(..., running=plan)` call instead of `_check_live_takeover`,
`_screen_preempted` and `_wifi_notice_pending` one after another. The
WiFi notice file is still read exactly where it was (the read is
throttled and deletes an expired file).
- A Vegas pass scans the live-priority plugins once instead of twice at the
same instant.
- The golden traces are byte-identical, and a capture of all 67 harness
runs in the suite (every sleep, frame, read and scan) matches `main`
apart from the duplicate scan above and one moment: in the 125 Hz loop a
live takeover's state change is made after the frame's 8 ms sleep rather
than before it, ending the screen at the same frame as before.
- New module: `src/screen_runner.py`. Core-internal: plugins have no reason
to import it, so it sets no `ledmatrix_min_version` floor.
### A scrolling screen held by its plugin's update() is reported
- While a plugin's `update()` runs it holds the plugin's lock, and that
plugin's frames are skipped: on a scroller, a frozen strip, with nothing
logged (and a freeze of 5 s or more is a gap, not a freeze, to the frame
stats). The high-FPS loop now times each run of skipped frames; one of
250 ms or more logs `Display of <plugin> held N ms by its update()`
(rate-limited per plugin) when it ends, and is recorded on the plugin's
health as a `display hold` busy skip, which never counts toward the
circuit breaker. The 1 Hz loop is left out: its frames are a second apart,
so one skipped frame there measures nothing and freezes nothing visible.
- Scroll speeds are solved against `limit_refresh_rate_hz`, so a panel that
cannot reach its cap ran every scroll slow by the shortfall, with no sign
why (one Pi 4 on a 120 Hz cap refreshed at ~110 Hz: 60 px/s ran at 55).
Once the display has measured the real rate over three windows of
scrolling, a panel more than 3% short of the cap is logged once, as a
warning from `src.common.frame_timing` that names a cap it can hold (a
multiple of 10, 5% under the measurement). The Display tab shows the same
under Limit Refresh Rate, with a button that fills it in, from the new
`GET /api/v3/config/refresh-rate`. Not checked in the emulator or on the
fallback canvas.
- The frame-stats file records `planned_refresh_hz` (additive), and the
scroll-speed advice behind the Vegas slider ignores a measurement written
under a different cap. Until now, after the cap changed, the slider kept
advising from the old rate until the display restarted.
- New in `src.common.scroll_config`: `refresh_shortfall()`, `holdable_cap()`
and `describe_refresh_shortfall()`.
### Fixed
@@ -122,18 +52,6 @@ Internal; no behaviour change. Stage 3 of `docs/RUN_LOOP_REDESIGN.md`.
changed frame, and the render loop writes it (`write_owed_snapshot()`)
once the interval has passed. The cadence is unchanged, and nothing extra
runs when no frame is owed.
- The installed-plugins list (`GET /api/v3/plugins/installed`) no longer
waits on GitHub. Its comment said the registry lookup made no network call,
but on a cold or expired cache `get_registry_info()` downloads plugins.json
(10 s timeout, three attempts), and with nothing cached to fall back on
every plugin's lookup repeated that: offline, 5 plugins took 11 s with DNS
failing and 2 plugins 65 s with the route black-holed, on every load. The
list now reads the registry copy already in memory, however old
(`get_cached_registry_info()`); with none yet it returns without update or
verified badges and starts one background refresh
(`refresh_registry_in_background()`, backing off for a minute after an
offline failure), so a later load has them. The store, install and update
paths still fetch as before.
### ESPN date-range fetches: fewer requests, fewer at once
@@ -158,18 +76,6 @@ soccer-scoreboard 2.39.2, alternating runs: **~450 requests per start, peak
spends one doomed 400 per window at every start (eleven at once from a
soccer board); the range is still retried `RANGE_RETRY_SECONDS` in.
### Fetch stats: bytes on the wire, not just decoded
`GET /api/v3/plugins/fetch-stats` reported only `bytes`, the decoded body
size, and that read as the download volume. ESPN gzips every scoreboard, so
it overstated what crossed the network about 14x: a college football
Saturday's scoreboard is 865 KB decoded and 63 KB on the wire, and ledpi's
"643 MB in 6 hours" of football was ~47 MB of actual traffic. Every counter
set (totals, per plugin, per host) now has `wire_bytes` too, read from
urllib3's count of the raw bytes it took off the socket. A response with no
urllib3 response behind it is counted at its decoded size. `bytes` keeps its
meaning.
### Cheap per-frame and per-fetch savings
- `BaseOddsManager.get_odds()` no longer pretty-prints every odds response
+3 -11
View File
@@ -649,9 +649,7 @@ When nothing is running on demand, `data.state` is
> on-demand machinery is internal — drive it through the REST endpoints
> above (or the web UI buttons). The API handlers
> (`start_on_demand_display()` / `stop_on_demand_display()` in
> `web_interface/blueprints/api_v3/display.py`) send the request over the
> display's control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)).
> Only when the socket cannot carry it do they write it into the cache
> `web_interface/blueprints/api_v3/display.py`) write a request into the cache
> manager under the `display_on_demand_request` key, which
> `DisplayController._poll_on_demand_requests()`
> (`src/display_controller.py`) picks up. A separate
@@ -749,14 +747,8 @@ keys helps troubleshoot stuck states.
"timestamp": 1234567890.123
}
```
**Purpose:** Communication from web interface to display controller, as the
fallback when the control socket cannot carry the request (deprecated; it
will be removed in a later release)
**When Set:** API endpoint receives a request and the display's control
socket is unavailable (display stopped, or older than the socket or the
command); some plugins also write it directly
**Read:** once a second while the display serves the control socket (0.25 s
without it), and only when the file changed since the last look
**Purpose:** Communication from web interface to display controller
**When Set:** API endpoint receives request
**Auto-Cleared:** After processing or 1 hour TTL
**2. display_on_demand_config** (No TTL)
+7 -12
View File
@@ -42,11 +42,11 @@ each other. They share three things:
| State | Where | Written by | Read by |
|---|---|---|---|
| On-demand command | control socket `/run/ledmatrix/control.sock` ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py), via [`src/ipc/client.py`](../src/ipc/client.py) | display: [`src/ipc/server.py`](../src/ipc/server.py) acks; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand request (fallback) | cache `display_on_demand_request` | web, only when the socket could not carry the request (`should_fall_back`); four plugins write it directly | display: `_poll_on_demand_requests()`, a `stat()` every 1 s while the socket is up (0.25 s without), read only when the file changed |
| On-demand request (fallback) | cache `display_on_demand_request` | web, when the socket fails; four plugins write it directly | display: `_poll_on_demand_requests()` |
| On-demand state | cache `display_on_demand_state` | display: `_publish_on_demand_state()` | web: `/api/v3/display/on-demand/status` |
| Current screen | cache `display_current_state` | display | web: `/api/v3/display/current-status` |
| Plugin errors | cache `plugin_error_snapshot` | display: `ErrorSnapshotPublisher` ([`src/error_aggregator.py`](../src/error_aggregator.py)) | web: `read_error_report()` for `/api/v3/errors/*` |
| Error clear | control socket `errors.clear`; cache `plugin_error_clear_request` as the fallback | web: `POST /api/v3/errors/clear` | display: applied before the socket answers; the mailbox on the error publisher's 5 s tick, read only when the file changed |
| Error clear | cache `plugin_error_clear_request` | web | display |
| Font usage | cache `font_usage_snapshot` | display: `FontUsagePublisher` ([`src/font_usage.py`](../src/font_usage.py)) | web: Fonts tab |
| Fetch statistics (requests per plugin and host) | cache `fetch_stats_snapshot` | display: `FetchStatsPublisher` ([`src/common/fetch_service.py`](../src/common/fetch_service.py)), at most once a minute on change | web: `read_fetch_stats()` for `/api/v3/plugins/fetch-stats` |
| Plugin health | cache `plugin_health:<id>` | display (web writes on reset) | web: `/api/v3/plugins/health` |
@@ -58,16 +58,11 @@ each other. They share three things:
The on-demand start route starts `ledmatrix.service` when it is not running
(`start_service`, on by default) but never restarts a running one. The routes
send the command over the display's control socket and get an ack; only when
the socket could not carry it (a stopped display, one older than the socket
or the command) do they write the mailbox instead. A display that had the
request and refused it is answered with the error, not posted a mailbox
copy. The display looks at the mailbox every
`MAILBOX_POLL_INTERVAL_WITH_SOCKET` (1 s) while it serves the socket, and
every `ON_DEMAND_POLL_INTERVAL` (0.25 s) without one, from its dwell sleep,
its render loops and Vegas's interrupt check as well as the main loop; a
look is one `stat()` unless the file changed. Both ways end in the same
handler, `_handle_on_demand_request()`.
send the command over the display's control socket and get an ack; when that
fails (a stopped display, one older than the socket) they write the mailbox
instead, which the display reads every `ON_DEMAND_POLL_INTERVAL` (0.25s), from
its dwell sleep, its render loops and Vegas's interrupt check as well as the
main loop. Both ways end in the same handler, `_handle_on_demand_request()`.
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
for the protocol, the permission model and the plan to retire the mailboxes.
+24 -114
View File
@@ -7,18 +7,14 @@ status. Stage 2 makes those commands land within a frame on every kind of
screen, and adds `brightness.set` and `plugin.reload`. Stage 3 adds a state
stream (`state.get`, `state.subscribe`), so the web interface reads what the
display is doing from the socket instead of from cache files the display
wrote to the SD card. Stage 4 makes the socket the only way a command goes
while it works: the web interface writes a mailbox only when the socket
cannot carry the request, `errors.clear` replaces the last command that
always went through a mailbox, and the display looks at the mailboxes once
a second, with a `stat()`. The file mailboxes and the cache keys stay as a
fallback for one release.
wrote to the SD card. The file mailbox and the cache keys stay as a fallback
for one release.
| | |
|---|---|
| Socket | `/run/ledmatrix/control.sock` (tmpfs) |
| Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` |
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop`, `POST /api/v3/plugins/update` (reload), `POST /api/v3/config/main` (brightness), `POST /api/v3/errors/clear`; and through [`web_interface/display_state.py`](../web_interface/display_state.py) (the state stream), `GET /api/v3/display/current-status`, `/display/on-demand/status`, `/plugins/installed` (`runtime`), `/plugins/state` and the reconciliations, `/health` (`display_loop`) |
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop`, `POST /api/v3/plugins/update` (reload), `POST /api/v3/config/main` (brightness); and through [`web_interface/display_state.py`](../web_interface/display_state.py) (the state stream), `GET /api/v3/display/current-status`, `/display/on-demand/status`, `/plugins/installed` (`runtime`), `/plugins/state` and the reconciliations, `/health` (`display_loop`) |
| Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it |
| Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it |
@@ -89,17 +85,6 @@ one. Clients branch on `error.code`, never on the message text.
| `plugin.reload` | `{plugin_id}` | `{plugin_id, reloaded: true, version, modes}` | queued, awaited (10 s) |
| `state.get` | `{since?, epoch?}` | a state snapshot (see "The state stream") | answered directly |
| `state.subscribe` | — | a state snapshot, then pushed `state` / `tick` events | answered directly, then a stream |
| `errors.clear` | `{cutoff: number}` (epoch seconds, finite, ≥ 0) | `{request_id, cutoff, cleared}` | answered directly, once applied |
`errors.clear` (stage 4) forgets the plugin errors the display recorded at
or before `cutoff` and rewrites its error snapshot (`plugin_error_snapshot`)
before it answers, so the web interface's next read already has it. The
request `id` is the clear's id, which the snapshot reports as
`applied_clear_id`. It is answered on the connection thread by a handler the
display registers (`ControlServer(handlers=...)`, the contract's
`DIRECT_COMMANDS`): the error aggregator and its publisher have their own
locks, and nothing the render thread owns is touched. A display that has no
handler answers `unknown_command`, as an older display does.
`duration` is a number of seconds, or a numeric string. `0`, `null` or `""`
mean "until stopped". `pinned` must be a real boolean: the REST route has
@@ -382,9 +367,7 @@ place it reads the mailbox:
Vegas iteration is on the stack (`_apply_pending_plugin_reloads`). Until
then the current screen ends early, as it does for a WiFi notice: the
frame loops, the dwell and Vegas's interrupt check all treat a pending
reload as a reason to stop (the frame loops through the Arbiter's
mid-screen check, `Source.RELOAD`; the dwell through
`_plugin_reload_pending`).
reload as a reason to stop (`_screen_preempted`).
- Only the quick half of the reload runs on the render thread
(`_start_plugin_reload`): the plugin's modes leave the rotation, its
config subscription is dropped, and `PluginManager.detach_plugin` takes
@@ -407,7 +390,7 @@ place it reads the mailbox:
unloads it mid-load; a disable saved meanwhile is applied once the
reload is done. A second reload of the same plugin runs after the first.
The floor on the mailbox read (0.25 s, 1 s since stage 4) does not apply to the queue, because
The 0.25 s floor on the mailbox read does not apply to the queue, because
draining it costs no disk read. A queued command also lets
`_service_pending_changes()` skip its own floor.
@@ -435,7 +418,7 @@ Now the queue wakes the render thread:
So a command lands within a millisecond or so on a static screen and in a
dwell, and within one frame in Vegas and on a scrolling screen. The mailbox
is slower on purpose (see "The mailboxes now"). Commands still run only on the render thread: the
keeps its old delays. Commands still run only on the render thread: the
connection threads only queue them and set the event. The one exception is
the slow half of `plugin.reload` (tearing down and loading the plugin),
which runs on its own thread. Every change to the display's state still
@@ -453,57 +436,11 @@ bookkeeping. A client's send to the render thread waking took 0.72 ms median
Without a socket (Windows, `LEDMATRIX_CONTROL_SOCKET=off`) the waits are the
plain sleeps they were.
**Exactly once.** Since stage 4 the web interface writes the mailbox only
when the display never had the request (see "When the web interface falls
back"), so a request goes one way or the other, never both. A command and a
mailbox write for the same request still share one `request_id`, and the
`on_demand_request_id` and processed-id checks still drop a second copy: an
older web interface (before stage 4) wrote the mailbox after a reply timed
out, too. The display takes such a copy out of the mailbox when it drops it.
## When the web interface falls back (stage 4)
The client tells a request the display never had from one it had and then
failed. `ControlError.sent` is True once the whole request was written to a
connected display; a refusal the display sends before reading anything
(`forbidden`, too many connections) carries no request id, and leaves it
False. `src.ipc.client.should_fall_back()` is the one rule every route uses:
| What happened | Example reasons | Mailbox? | The route answers |
|---|---|---|---|
| The display never had it | `no_socket`, `refused`, `disabled`, `unsupported`, a connect or send that timed out, `forbidden` / `busy` at the door, `invalid_request` (refused by the client itself) | yes | success, `transport: "mailbox"`, `socket_error` |
| A display too old to know it (the upgrade case) | `unknown_command`, `unsupported_version` | yes | as above |
| The display had it and failed | `busy` (queue full), `invalid_args`, `internal`, a timeout or hang-up after the send, `bad_response` | no | `503` (`400` for `invalid_args`), `socket_error` |
A display that had the request may have applied it (a reply that timed out),
or would refuse the mailbox copy as well (bad arguments), or is stuck and
would not read the mailbox either (a full queue). Writing the copy anyway
only turned that into a "success". An on-demand stop with `stop_service`
still stops the service, which ends on-demand whatever happened.
Brightness and plugin reload never had a mailbox: without the socket, the
config watcher applies the saved brightness and a reload becomes the
restart banner, as before.
### The mailboxes now
| Mailbox | Written by | Read by the display | While the socket is up |
|---|---|---|---|
| `display_on_demand_request` | the web interface, only on fallback; four plugins directly (birdnet-go, mqtt-notifications, on-air, pomodoro-timer) | the render thread, `_poll_on_demand_requests()` | looked at every 1 s (`MAILBOX_POLL_INTERVAL_WITH_SOCKET`), 0.25 s without a socket |
| `plugin_error_clear_request` | the web interface, only on fallback | the error publisher's thread, every 5 s tick | unchanged rate |
A look is one `stat()` of the mailbox file (`CacheManager.file_signature`):
`(inode, mtime, size)`, and every write renames a new file into place, so a
new write always looks different. `MailboxWatch` reads the file only when
that changed since the last look, so a mailbox that holds nothing new, or
nothing at all, costs no open and no parse. A socket command never reads or
deletes the on-demand mailbox. A start already processed is taken out of
the mailbox instead of being re-read until it expires.
A request that comes through the on-demand mailbox while the socket is up
is logged once per writer (`came through the file mailbox although the
control socket is up`), which names the plugins that still need an
in-process way in before the mailbox is removed.
**Exactly once.** A command and a mailbox write for the same request share
one `request_id`. If the client times out after the display queued the
command and then also writes the mailbox, the display processes the request
once. The existing `on_demand_request_id` and processed-id checks drop the
second copy.
## Robustness
@@ -520,10 +457,9 @@ block the render loop or crash it:
and the connection is closed, because the next message boundary cannot be
found. A client that disconnects mid-message is dropped silently. No
exception from a handler leaves the connection thread.
- **Full queue.** When the queue is full, the client gets `busy`, and the web
interface answers `503` rather than write the mailbox, which the stuck
render thread would not read either. A full queue means the render thread
is stuck, and the systemd watchdog deals with that.
- **Full queue.** When the queue is full, the client gets `busy` and falls
back to the mailbox. A full queue means the render thread is stuck, and the
systemd watchdog deals with that.
- **Awaited commands.** The wait for an awaited command's outcome happens on
its connection thread and is bounded (`AWAIT_SECONDS`), so a stuck render
thread costs that client `pending` and one connection slot for at most
@@ -640,30 +576,15 @@ device never touches the live display.
an uninstall that keeps its config, still answer `restart_required`.
They can now use a load/unload command and report the result the same
way the update route does.
4. **The mailboxes become a fallback (done).**
- The web interface writes a mailbox only when the socket could not carry
the request (`should_fall_back`); a display that had it and failed is
answered as that (see "When the web interface falls back").
- `errors.clear` replaces `plugin_error_clear_request` as the way a clear
reaches the display.
- The display looks at the on-demand mailbox once a second while the
socket is up, reads either mailbox only when its file changed, and logs
who still writes the on-demand one (see "The mailboxes now").
- Not changed, deliberately: config saves (the schedule, the dim
schedule, plugin settings) still reach the display through
`config.json` and its watcher, which is the setting itself rather than
a message; see `config.reload` under stage 2. The preview viewer marker
(`/tmp/led_matrix_preview_viewer`) is a presence signal the display
already stats at most once a second. Plugin health and metrics resets
write the persisted record the display publishes and do not reach the
running display (their routes say so); they are not mailboxes.
5. **Remove the mailboxes (next release).** Once every device has run a
display with stage 4, the web interface stops writing both mailboxes and
the display stops reading them. The four plugins that write
`display_on_demand_request` need an in-process way to ask for the screen
first. The display also stops writing `display_current_state`,
`display_on_demand_state` and `plugin_runtime_snapshot` once the web
interface no longer falls back to them.
4. **Retire the mailboxes.** After a release in which every device has had the
socket, the web interface stops writing `display_on_demand_request`, and
the display stops polling it, logging the plugins that still write it so
they can move to an in-process `request_display()`. The other cache keys
used as messages (`plugin_error_clear_request` and the remaining
`display_*` keys) move to the socket or to tmpfs. The display also stops
writing `display_current_state`, `display_on_demand_state` and
`plugin_runtime_snapshot` once the web interface no longer falls back to
them.
## Checking it on a device
@@ -678,18 +599,7 @@ curl -s -X POST localhost:5000/api/v3/display/on-demand/start \
If the response says `"transport": "mailbox"`, `socket_error` gives the
reason. `no_socket` means the display is stopped or predates the socket.
`refused` usually means the web user is not in the socket's group, which
takes effect when the web service restarts after the user is added. A `503`
with `"transport": "socket"` means the display had the request and did not
take it (`busy`, `timeout`, ...): nothing was written to the mailbox.
An error clear:
```bash
curl -s -X POST localhost:5000/api/v3/errors/clear \
-H 'Content-Type: application/json' -d '{"all":true}'
# ... "applied": true, "transport": "socket"
sudo journalctl -u ledmatrix | grep -E "Cleared .* plugin error|file mailbox"
```
takes effect when the web service restarts after the user is added.
Brightness and a plugin reload:
+16 -36
View File
@@ -464,7 +464,7 @@ Request a specific plugin to display on-demand.
- `mode` (string, optional): Display mode name (plugin_id inferred if not provided)
- `duration` (number, optional): Duration in seconds (0 = until stopped)
- `pinned` (boolean, optional): Pin display (pause rotation)
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within a frame over its control socket (within about a second through the mailbox fallback). When false and the service is stopped, the route returns 400.
- `start_service` (boolean, optional): Start the display service if it is not running (default: true). A running service is never restarted: it picks the request up within about a quarter of a second. When false and the service is stopped, the route returns 400.
**Response**:
```json
@@ -489,19 +489,10 @@ display's control socket acknowledged it (it is queued for the render thread,
which wakes for it and applies it within a frame; see
[IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), `"mailbox"` means it was
written to the cache mailbox the display polls, as before the socket existed.
The mailbox is used only when the socket could not carry the request. With
`"mailbox"`, `socket_error` gives the reason (`no_socket` when the display is
stopped or predates the socket, `refused`, a connect `timeout`,
`unknown_command` from a display too old for the command, ...). Either way
the request is applied the same way; `request_id` is the same id in both.
When the display had the request and did not take it -- a full queue
(`busy`), bad arguments (`invalid_args`), no answer after the request was
sent (`timeout`, `closed`) -- the route answers `503` (`400` for
`invalid_args`) with `status: "error"` and `data: {request_id, transport:
"socket", socket_error}`, and writes nothing to the mailbox. The stop route
does the same, except that with `stop_service: true` it still stops the
service and answers success.
With `"mailbox"`, `socket_error` gives the reason the socket was not used
(`no_socket` when the display is stopped or predates the socket, `timeout`,
`refused`, `busy`, ...). Either way the request is applied the same way;
`request_id` is the same id in both.
### Stop On-Demand Display
@@ -2257,11 +2248,13 @@ error with `"all": true` (`max_age_hours` is then ignored).
}
```
The clear goes to the display service over its control socket
(`errors.clear`, see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), which
applies it, rebuilding its counts from the errors it keeps, and republishes
before it answers: `applied` is `true`, `transport` is `"socket"`, and
`cleared_count` is the display's own count.
The clear is asynchronous. The web interface records a request
(`plugin_error_clear_request` in the shared cache), and the display service
applies it within about 5 seconds, rebuilding its counts from the errors it
keeps and republishing. Reads hide the cleared errors from the moment the
request is recorded. Until the display service applies an age-based clear,
`recent_errors` and `active_patterns` are already filtered but the counts
are the old ones, and `clear_pending` is `true`.
```json
{
@@ -2269,30 +2262,17 @@ before it answers: `applied` is `true`, `transport` is `"socket"`, and
"data": {
"cleared_count": 13,
"clear_requested": true,
"applied": true,
"transport": "socket",
"request_id": "5f0c1e...",
"cutoff": "2026-09-23T09:59:02.310000"
},
"message": "Cleared all errors"
"message": "Clear of all errors requested; the display service applies it within about 5 seconds"
}
```
When the socket cannot carry it (the display is stopped, or older than
`errors.clear`) the clear is asynchronous, as before: the web interface
records a request (`plugin_error_clear_request` in the shared cache),
`applied` is `false` and `transport` is `"mailbox"`, and the display service
applies it within about 5 seconds. Reads hide the cleared errors from the
moment the request is recorded. Until the display service applies an
age-based clear, `recent_errors` and `active_patterns` are already filtered
but the counts are the old ones, and `clear_pending` is `true`. Then
`cleared_count` is how many of the reported errors the clear hides, and
`cleared_count` is how many of the reported errors the clear hides. It is
`null` when that cannot be known before the display service applies it (an
age-based clear over more errors than the report lists).
A request that could not be written to the shared cache answers `500`. A
display that had the request and failed it (`internal`, a timeout after the
request was sent) answers `503`, with `context.socket_error`.
age-based clear over more errors than the report lists). A request that
could not be written to the shared cache answers `500`.
---
+81 -205
View File
@@ -36,189 +36,122 @@ Each pass, in order:
1. `loop_pass()` (watchdog). Apply a pending plugin enable/disable, then
any plugin reloads the control socket asked for
(`_apply_pending_plugin_reloads`; a pending reload ends the screen
before it, like a WiFi notice, as `Source.RELOAD` at the runner's
service points). The static screen's frame sleep and the dwell wait on
the socket's queue instead of sleeping (`_wait_frame_interval`,
`_sleep_with_plugin_updates`); without a socket, as in the golden
traces, they are the plain sleeps.
before it, like a WiFi notice, through `_screen_preempted`). The static
screen's frame sleep and the dwell wait on the socket's queue instead of
sleeping (`_wait_frame_interval`, `_sleep_with_plugin_updates`); without
a socket, as in the golden traces, they are the plain sleeps.
2. With no modes: dwell 1 s, next pass.
3. Poll on-demand requests and expiry, release plugins loaded only for
on-demand, tick plugin updates, drop an expired WiFi notice, evaluate
the schedule (an on-demand session overrides scheduled-off), apply the
brightness target. Then gather the Arbiter's inputs
(`_arbiter_inputs`) and call `Arbiter.decide()`.
(`_arbiter_inputs`) and call `Arbiter.decide()`, which picks one of
steps 4-6 or returns `LEGACY` for steps 7-9 (stage 2).
4. **Scheduled off:** blank, dwell up to 60 s. `_blank_while_scheduled_off`
5. **Follower:** render one frame from the leader. `_run_follower_frame`
6. **WiFi notice** (unless on-demand): draw it, dwell 0.5 s. `_show_wifi_notice`.
It also ends a running screen within about a second (the runner's
service points), and a screen cut short resumes after it.
7. **The Sources below the notice:** read whether Vegas is on and make the
live-priority scan (`_arbiter_inputs_below_wifi`, where run() always
read them), and call `decide()` again. It answers OnDemand (the
session's current mode), Live (the next live mode, round-robin; a game
that goes live during a screen takes over at the next service point, at
most once a second), Vegas (`LEGACY`) or Rotation. `_take_plan` applies
the answer: a live claim or the resume when live priority ends, the
on-demand index.
8. **Vegas** (`_run_vegas_iteration`): one iteration of up to
`max_cycle_duration`. A completed iteration ends the pass, and so does
one that yielded for the schedule, a reload or a WiFi notice. Any other
interrupted one asks `decide()` once more (`vegas_yielded`): a game that
stopped the ticker, or an on-demand session that started, shows next.
9. **One screen:** pick the plugin (`_plugin_for_mode`) and hand the plan
to the `ScreenRunner` (`src/screen_runner.py`). It draws the first frame
through the executor (`_dispatch_first_frame`), has the controller fill
in the plugin's durations, dynamic flag and frame policy
(`_complete_plan`), runs the 125 Hz or 1 Hz frame loop with a service
point after each frame, makes up the minimum duration, and returns an
`Outcome`. On `PREEMPTED` the pass ends without advancing. On no
content, rotate at once (`_note_empty_pass`, `_skip_failed_plugin_modes`).
Otherwise `ArbiterState.after()` picks the next mode
(`_advance_after_screen`).
It is also polled mid-screen (`_wifi_notice_pending`): the frame loops,
the dwell sleep and an interrupted Vegas iteration end within about a
second when one arrives, and a screen cut short resumes after it.
7. **Live priority** (unless on-demand, or Vegas keeps live content in the
ticker): switch to the next live mode, or resume the rotation. A game
that goes live during a screen is caught sooner, by
`_check_live_takeover` in the frame loops and the dwell sleep (at most
once a second, and not while a live mode is showing).
8. **Vegas** (unless on-demand, or live content preempts it): run one
iteration of up to `max_cycle_duration`. A completed iteration ends the
pass, and so does one that yielded for a WiFi notice or the schedule.
Any other interrupted one falls through to step 9 in the same pass.
9. **One screen:** pick the mode (`_resolve_active_mode`), the plugin
(`_plugin_for_mode`), draw the first frame through the executor
(`_dispatch_first_frame`). On no content, rotate at once
(`_note_empty_pass`, `_skip_failed_plugin_modes`). Otherwise work out the
bounds (`_track_dynamic_cycle`, `_resolve_durations`,
`_clamp_to_on_demand`) and the frame rate (`_needs_high_fps`), run the
125 Hz or 1 Hz frame loop, make up the minimum duration, then pick the
next mode (`_advance_after_screen`).
## Design
The helpers named above were extracted in stage 1 without changing
behaviour. Since stage 2 the choice between steps 4, 5, 6 and the rest is
made by `Arbiter.decide()` in `src/display_arbiter.py`. The frame loops, the
Vegas branch and every early exit are still inline in `run()`.
## Target design
```python
def run(self):
while True:
inputs = self._arbiter_inputs() # schedule, follower, notice
plan = Arbiter.decide(self._arbiter_state(), inputs, now)
... # off / follower / notice
plan = self._take_plan(Arbiter.decide(state, self._arbiter_inputs_below_wifi(inputs), now))
if plan.source is Source.LEGACY: # Vegas, until stage 4
plan = self._run_vegas_iteration(...)
outcome = runner.run(plan, plugin) # ExitReason + elapsed
if outcome.exit_reason is not ExitReason.PREEMPTED:
self._advance_after_screen(plan, outcome) # ArbiterState.after
inputs = self._drain_inputs() # requests, schedule, config, sync
plan = self.arbiter.decide(self.state, inputs, clock.now())
outcome = self.runner.run(plan) # ExitReason + elapsed
self.state = self.state.after(plan, outcome) # rotation, on-demand index, live resume
```
The controller's attributes (`current_display_mode`, `current_mode_index`,
`on_demand_*`, `_live_resume_index`) stay the record that the web UI, the
control socket and the on-demand cache read. `_arbiter_state()` snapshots
them into a frozen `ArbiterState`; the transitions are pure methods on it,
and the controller writes their result back (`_adopt_state`).
### Sources
Each kind of content is a Source. A Source looks at the state and the
inputs and either offers a screen or passes. The Arbiter asks them in this
order:
| Order | Source | Offers a screen when | Code |
| Order | Source | Offers a screen when | Today |
|---|---|---|---|
| gate | ScheduledOff | the schedule is off and no on-demand session overrides it | `decide` |
| 1 | Follower | a sync leader is driving this panel | `decide` |
| 2 | OnDemand | a session is active (its mode list, index, expiry and pin) | `_on_demand_plan` |
| 3 | Wifi | a status message is pending and on-demand is not active | `decide` |
| 4 | Live | a live-priority plugin has live content (round-robin across several) | `live_pick` |
| 5 | Vegas | Vegas is enabled and nothing above wants the panel | `LEGACY`, run by `_run_vegas_iteration` |
| 6 | Rotation | always: the rotation's current mode | `rotation_plan` |
| gate | ScheduledOff | the schedule is off and no on-demand session overrides it | step 4 |
| 1 | Follower | a sync leader is driving this panel | step 5 |
| 2 | OnDemand | a session is active (its mode list, index, expiry and pin) | `_resolve_active_mode` |
| 3 | Wifi | a status message is pending and on-demand is not active | step 6 |
| 4 | Live | a live-priority plugin has live content (round-robin across several) | step 7 |
| 5 | Vegas | Vegas is enabled and nothing above wants the panel | step 8 |
| 6 | Rotation | always: `available_modes[current_mode_index]` | step 9 |
ScheduledOff is a gate in front of the Sources because that is how it works
today: a scheduled-off panel stays blank even for a follower, and only an
on-demand session overrides it.
The Rotation answers `state.current_mode`, not
`available_modes[current_mode_index]`: the two agree except where something
moved the panel off the list and the rotation carries on from there (a live
mode no rotation entry names, or None after a session ended with no enabled
mode to resume to), and `run()` always showed `current_display_mode`.
### Arbiter
```python
Arbiter.decide(state, inputs, now, running=None) -> ScreenPlan
Arbiter.decide(state, inputs, now) -> ScreenPlan
```
`decide` is a pure function: it does no I/O, takes no locks and does not
sleep. It can be tested with plain tables of (state, inputs, now) mapped to
an expected plan.
- `ArbiterState`: the current mode; the rotation and its index; the
on-demand session's modes, index, expiry and pin; the live resume point;
whether a mid-screen takeover has not shown yet. Transitions:
`next_on_demand`, `showing`, `claim_live`, `release_live`, `after`.
- `ArbiterInputs`: whether the schedule has the panel on, an on-demand
session, a follower, the WiFi notice, the live modes (None where no scan
was made), whether Vegas is on and keeps live content in its ticker,
whether this pass's Vegas iteration has yielded, and (mid-screen) whether
a plugin reload is waiting.
- `ScreenPlan`:
an expected plan. It returns a `ScreenPlan`:
| Field | Meaning |
|---|---|
| `source` | which Source won |
| `mode`, `plugin` | what to draw (None for a blank or follower plan); the plugin id once resolved |
| `min_duration`, `max_duration` | from `_resolve_durations` and the on-demand bound (`on_demand_bound`), filled in after the first frame; an on-demand plan's `max_duration` is what is left of the session at `now` |
| `mode`, `plugin` | what to draw (None for a blank or follower plan) |
| `min_duration`, `max_duration` | from `_resolve_durations` and `_clamp_to_on_demand` |
| `dynamic` | run until the plugin's cycle completes, between min and max |
| `frame_policy` | `HIGH_FPS` or `STATIC`, today `_needs_high_fps`; see stage 5 |
| `frame_policy` | today `_needs_high_fps` (125 Hz or 1 Hz); see stage 5 |
| `preemptible_by` | the Sources allowed to interrupt this plan mid-screen |
| `notice`, `deadline`, `ends_live` | the WiFi notice; the on-demand expiry for the bound; "live priority just ended, resume the rotation first" |
`decide` cannot ask a plugin anything, so the fields a plugin answers are
filled in by the controller after the first frame, where they were always
read (`_complete_plan`).
With `running`, `decide` answers the mid-screen question instead: `running`
itself while the screen holds, else the plan that ends it
(`_hold_or_preempt`), in the order the frame loops always checked:
1. Live: a game went live while a non-live screen runs. It is the one
preemption that changes the state (the rotation moves to the live mode
and remembers where it was), and it is claimed even when a WiFi notice
is also pending; the next pass shows the notice, then the game.
2. The panel's mode moved under the screen (on-demand started, ended or
changed mode; the rotation was rebuilt).
3. The schedule turned the panel off.
4. A WiFi notice (unless on-demand outranks it), compared with its expiry.
5. A plugin reload is waiting (between frames only).
Every screen is preemptible by the gate, OnDemand, Wifi, Live, Rotation and
a reload (`SCREEN_PREEMPTERS`), except that a live screen leaves Live out
(`LIVE_PREEMPTERS`): live games take turns between screens. A follower and
Vegas are looked at only between screens.
### ScreenRunner
```python
ScreenRunner(clock: FrameClock, host: ScreenHost).run(plan, plugin) -> Outcome
ScreenRunner(clock: FrameClock).run(plan) -> Outcome(exit_reason, elapsed)
```
The ScreenRunner draws the first frame (`_dispatch_first_frame`), runs the
frame loop that the plan's frame policy selects, services pending changes
between frames, and returns one `ExitReason`:
| ExitReason | Golden-trace exit |
| ExitReason | Today's equivalent (golden-trace exit) |
|---|---|
| `DURATION` | target duration reached (`duration`) |
| `CYCLE_COMPLETE` | dynamic plugin finished after its minimum (`cycle-complete`) |
| `EMPTY` | first frame returned False, or no plugin (`empty`; `raised` when display() raised inside the executor; `no-plugin`, `breaker`) |
| `EMPTY` | first frame returned False (`empty`; `raised` when display() raised inside the executor) |
| `ERROR` | the dispatch itself raised (`error`) |
| `DISPLAY_FALSE` | a later frame returned False (`display-false`) |
| `PREEMPTED` | another Source took the panel (`on-demand-*`, `schedule-off`, `live`, `wifi`, ...) |
| `RELOAD` | a plugin reload is waiting: the screen ends early but counts as shown, and the rotation advances |
| `PREEMPTED` | another Source took the panel (`on-demand-*`, `schedule-off`, `vegas-interrupt`, ...) |
`PREEMPTED` replaces the five `current_display_mode != active_mode` checks.
The runner asks its host at named service points (`Checkpoint`): `FRAME`
after each frame (and when a socket command wakes the 1 Hz wait),
`AFTER_LOOP` / `AFTER_COMPLETED_LOOP` when the frame loop ends,
`after_dwell` after the make-up dwell, and `FINAL` before the rotation
advances. Each is one `decide(..., running=plan)` call
(`DisplayController._screen_check`). The checkpoint says whether a pending
reload counts there and when the WiFi notice file is read (`NoticeRead`):
the read is throttled to once a second and deletes an expired file, so it
happens exactly where the loop always read it.
The runner asks the Arbiter, at the throttled service points it already has,
whether a Source in `plan.preemptible_by` now wants the panel.
In the 125 Hz loop the live-priority scan is made before the frame's sleep
(`_screen_service`), at the moments it always was, and weighed by the
service point after the sleep, where the loop always decided to end the
screen.
`FrameClock` provides `time()`, `perf_counter()` and `sleep()`, the shape of
the `time` module. In production it is `_ModuleClock`, which looks up
`src.display_controller.time` on each call, so the golden traces' fake clock
drives the runner as it drove the inline loops. The runner's log lines use
the controller's logger, so they keep their source in the journal.
`FrameClock` provides `now()` and `sleep()`. In production it is
`time.monotonic`/`time.sleep`. In the golden traces it is the fake clock
that the harness patches in today.
## Stages
@@ -317,84 +250,37 @@ What shipped:
Source, the dwells, the expiry comparison, the snapshot's reads, each
dispatch in `run()`); every one failed a test.
### Stage 3: ScreenRunner and `PREEMPTED` (done; awaiting the ledpi soak)
### Stage 3: ScreenRunner and `PREEMPTED`
The plan, from where stage 2 left off:
Move the two frame loops, the make-up dwell and the dynamic-duration exit
into `ScreenRunner.run(plan)` with an injected `FrameClock`. Replace the
five re-checks with `PREEMPTED`. Add the OnDemand, Live and Rotation Sources
so `LEGACY` is left meaning only Vegas.
Concretely, from where stage 2 left off:
1. `ArbiterState` gains the rotation index, the on-demand mode list, index,
expiry and pin, and the live resume point. `ArbiterInputs` gains the
live modes and whether Vegas is enabled and keeps live content in the
ticker.
2. OnDemand returns its current mode with the session's bound, reading
`now` for the expiry. Live returns the next live mode (round-robin).
Rotation returns the rotation's mode. `ScreenPlan` gains `mode`,
`plugin`, `min_duration`, `max_duration`, `dynamic`, `frame_policy` and
`preemptible_by`.
3. `ScreenRunner.run(plan)` returns an `ExitReason`; `state.after(outcome)`
replaces `_advance_after_screen`'s step and the live-resume bookkeeping.
Each mid-screen check asks `decide()` whether a Source in
`plan.preemptible_by` now wins.
expiry and pin, and the live resume point (today `current_mode_index`,
`on_demand_*` and the live-priority stash). `ArbiterInputs` gains the
live modes (`_collect_live_modes`) and whether Vegas is enabled and keeps
live content in the ticker.
2. OnDemand returns its current mode with `_clamp_to_on_demand`'s bound,
reading `now` for the expiry. Live returns the next live mode
(round-robin). Rotation returns `available_modes[current_mode_index]`.
`ScreenPlan` gains `mode`, `plugin`, `min_duration`, `max_duration`,
`dynamic`, `frame_policy` and `preemptible_by`.
3. `ScreenRunner.run(plan)` returns an `ExitReason`; `state.after(plan,
outcome)` replaces `_advance_after_screen` and the live-resume
bookkeeping. Each mid-screen check asks `decide()` whether a Source in
`plan.preemptible_by` now wins, so `_screen_preempted`,
`_check_live_takeover` and `_wifi_notice_pending` become one call.
4. The control socket (`_drain_control_commands`, `_wait_for_control`) and
state publishing stay where they are; the runner calls them at its
service points.
What shipped, one commit each: the runner; then the OnDemand, Live and
Rotation Sources; then one `decide()` call at the service points.
- `src/screen_runner.py` (on the mypy ratchet): `ScreenRunner`,
`FrameClock`, `ExitReason`, `Outcome`, `Checkpoint`, `NoticeRead`,
`Screen` and the `ScreenHost` protocol, which `DisplayController`
implements through `_ScreenHost` (one-line forwards to its own methods).
The two frame loops, the make-up dwell and the dynamic-duration exit
moved in unchanged, pacing included.
- `src/display_arbiter.py`: `Source` gains `ON_DEMAND`, `LIVE`, `ROTATION`
and `RELOAD`; `LEGACY` means only Vegas. `FramePolicy`. `ArbiterState`
and `ArbiterInputs` as listed under "Arbiter". The pure helpers
`on_demand_bound` (`_clamp_to_on_demand`), `live_pick`
(`_check_live_priority`'s pick), `live_takeover` (the mid-screen claim)
and `rotation_plan`.
- A pass asks `decide()` twice: once with the inputs every pass reads, and
once, only when nothing above the notice took the panel, with the Vegas
check and the live scan, read where run() always read them (the scan
asks every live-priority plugin, and the Vegas check applies queued
Vegas config, so reading them earlier, on a follower or notice pass,
would be a change). A Vegas iteration that yields asks a third time.
- `_resolve_active_mode`, `_clamp_to_on_demand` and `_screen_preempted` are
gone. `_apply_live_priority`, `_check_live_priority`,
`_check_live_takeover` and `_wifi_notice_pending` remain (Vegas, the
dwell sleep and the tests call them), built on the same pure rules.
- `_sleep_with_plugin_updates` keeps its own break rules. It also serves
the blank, the notice and the idle wait, which are not screens, and its
rules are edge-triggered (an on-demand session starting on the mode
already showing ends a dwell but not a frame loop); folding them into
`decide()` would change behaviour.
Behaviour, checked three ways:
- Golden traces: unchanged, no regeneration.
- Every harness run in the suite (67: the goldens plus the live-takeover,
WiFi+live, socket-wake, plugin-reload, schedule and tick tests) was
captured with every sleep, `display()` call, WiFi read, live scan,
publish, dwell and scroll-state call logged, and diffed against
`origin/main`. Identical, except:
- a Vegas pass used to scan the live plugins twice at the same instant
(step 7, then step 8's "is anything live?"); it scans once;
- `_apply_live_priority(None)` calls that changed nothing are not made;
- throttled WiFi reads that returned the cached answer (no side effect)
after a notice had already ended the screen are not made;
- in the 125 Hz loop the live scan still runs before the frame's sleep,
but the claim is made by the service point after it, so the "live"
state change happens 8 ms later. The screen ends at the same frame as
before.
- Tables: `test/test_display_arbiter.py` (OnDemand, Live, Vegas/Rotation,
`after`, the 24-row mid-screen table, `live_takeover`) and
`test/test_screen_runner.py` (the runner on a scripted host and fake
clock; the controller's service point and the reads it makes). A
mutation run broke each moved or new piece once; see the PR.
This stage touches frame pacing (the 8 ms deadline sleep, the 1 ms yield),
so it needs a frame soak on ledpi, A/B against main, before it merges.
Coordinate with whoever owns scroll performance (`docs/SCROLL_PERFORMANCE.md`).
so it needs a frame soak on ledpi, A/B against main. Coordinate with
whoever owns scroll performance (`docs/SCROLL_PERFORMANCE.md`).
### Stage 4: Vegas as a Source
@@ -459,13 +345,3 @@ PR that updates the affected trace and explains why. All six are fixed:
A new one found later goes the same way: record it here with the trace that
shows it, then fix it in its own PR, not inside a restructure stage.
Open:
- Vegas stops for a sync follower (its interrupt check includes
`is_follower_active`), but the yield path never looks at a follower, so a
full rotation screen (20 s in the test) runs before the next pass hands
the panel to the leader. Found by stage 3's mutation run;
`test_screen_runner.py::TestThroughRun::test_vegas_yielding_to_a_follower_shows_a_rotation_screen_first`
pins it. Stage 4, which drops the interrupt callback, is the natural
place to fix it.
+25
View File
@@ -77,6 +77,31 @@ The Vegas **Scroll Speed** slider in the web UI shows the same thing live: a
line under it says what your speed will run as on this panel, and links to the
nearest smooth speeds.
### A panel that cannot reach its cap
Speeds are solved against `limit_refresh_rate_hz`, the configured cap, but a
cap is only a ceiling: a long chain, a high `pwm_bits` or a big
`gpio_slowdown` can leave the panel below it. One Pi 4 driving 2×128×64 on
`adafruit-hat-pwm` with `pwm_bits 9` and `gpio_slowdown 5` measured
107.6–113.1 Hz under a 120 Hz cap. Frames still move whole pixels, but
every scroll runs that much slower than configured (60 px/s ran at 55 px/s),
and the smooth speeds are the cap's rather than the panel's.
The display measures the real rate from its own frames. About a minute
into scrolling, a panel more than 3% short of its cap is logged once:
```
WARNING - src.common.frame_timing - The panel refreshes at about 113 Hz, below
the 120 Hz that scroll speeds are planned for ... Set Limit Refresh Rate to
100 Hz (web UI, Display tab), which this panel can hold, and restart.
```
The Display tab says the same under **Limit Refresh Rate**, with a button
that fills in the suggested cap (`GET /api/v3/config/refresh-rate`). The
suggestion is a multiple of 10 at least 5% under the measurement, because
an uncapped panel drifts and the measurement is the fast end of it. A cap the
panel holds also stops the drift.
### How a slow speed stays crisp
`SwapOnVSync(canvas, framerate_fraction)` holds each frame for N panel
-1
View File
@@ -88,7 +88,6 @@ src/plugin_system/testing/vegas.py
src/plugin_system/vegas_elements.py
src/redaction.py
src/scan_order.py
src/screen_runner.py
src/startup_validator.py
src/vegas_mode/__init__.py
src/vegas_mode/config.py
-6
View File
@@ -14,12 +14,6 @@ project_dir = os.path.dirname(os.path.abspath(__file__))
if project_dir not in sys.path:
sys.path.insert(0, project_dir)
# Cap glibc's malloc arenas before any thread exists (arenas already made
# stay): the in-process twin of the unit's MALLOC_ARENA_MAX=2, for units
# installed before that line. A no-op off glibc. See src/malloc_tuning.py.
from src import malloc_tuning
print('EXP cap_arenas=%s' % malloc_tuning.cap_arenas(), flush=True)
# Under systemd the watchdog clock is already running, and start-up (plugin
# loads, initial updates) takes far longer than the render loop's limit. Widen
# it before anything slow is imported; the render loop narrows it again once
+2 -57
View File
@@ -28,7 +28,7 @@ import os
import time
from datetime import datetime
import pytz
from typing import Any, Dict, List, Optional, Tuple
from typing import Any, Dict, List, Optional
import logging
import threading
import tempfile
@@ -72,43 +72,6 @@ def _outlived(record: Any, max_age: Optional[float], now: float) -> bool:
return False
_NOT_SEEN: Any = object()
class MailboxWatch:
"""Tells the poller of a mailbox key whether its file changed since the
last look, from one stat() (:meth:`CacheManager.file_signature`).
The display polls the mailboxes the web interface falls back to. Reading
one is an open and a JSON parse; with this a poll that finds the same file
(or none) costs a stat, and the file is read only after a new write. A
cache without ``file_signature`` (a test double) is read every time.
"""
def __init__(self, key: str):
self.key = key
self._seen: Any = _NOT_SEEN
def changed(self, cache_manager: Any) -> bool:
"""True when the poller should read the key now."""
signature = getattr(cache_manager, 'file_signature', None)
sig = signature(self.key) if callable(signature) else _NOT_SEEN
if sig is not None and not isinstance(sig, tuple):
return True # cannot tell: read it
if sig is None:
self._seen = None
return False # no file, nothing to read
if sig == self._seen:
return False
self._seen = sig
return True
def forget(self) -> None:
"""Read the key on the next poll even if its file has not changed
(the last read failed)."""
self._seen = _NOT_SEEN
class CacheManager:
"""Manages caching of API responses to reduce API calls."""
@@ -332,25 +295,7 @@ class CacheManager:
def _get_cache_path(self, key: str) -> Optional[str]:
"""Get the path for a cache file."""
return self._disk_cache_component.get_cache_path(key)
def file_signature(self, key: str) -> Optional[Tuple[int, int, int]]:
"""``(st_ino, st_mtime_ns, st_size)`` of ``key``'s file, or None when
there is no file (the key is absent, or this cache has no disk tier).
One stat(), no read: a poller of a mailbox another process writes
compares it with the last one it saw and reads the file only when it
changed. Every write replaces the file (a temp file renamed into
place), so a new write always has a new inode, however fast it came.
"""
path = self._get_cache_path(key)
if not path:
return None
try:
st = os.stat(path)
except OSError:
return None
return (st.st_ino, st.st_mtime_ns, st.st_size)
def get_cached_data(self, key: str, max_age: int = 300, memory_ttl: Optional[int] = None) -> Optional[Dict[str, Any]]:
"""Get data from cache (memory first, then disk) honoring TTLs.
+1 -30
View File
@@ -59,10 +59,7 @@ says how old with ``cache_max_age`` (``fetch_get(..., cache_max_age=ttl)``;
Identical means what the validator store keys on: URL, query, effective
headers and, for a session with cookies or auth, the session.
**Counters.** Requests, merged requests, bytes (``bytes`` decoded, as the
caller reads them; ``wire_bytes`` as they crossed the network, which is
what a metered connection pays for -- ESPN gzips, so the two differ ~14x),
304s, errors, HTTP errors,
**Counters.** Requests, merged requests, bytes, 304s, errors, HTTP errors,
adapter retries, throttled requests and seconds waited, plus requests
answered without the network: ``memo_hits`` (the response cache) and
``cache_hits`` / ``legacy_cache_hits`` (a shared ESPN scoreboard cache entry,
@@ -204,7 +201,6 @@ _COUNTER_FIELDS = (
"throttled", # requests that waited for a host budget
"overruns", # requests that went after max_wait_seconds anyway
"bytes", # decoded response body bytes received
"wire_bytes", # body bytes as they came off the socket (still compressed)
"wait_seconds", # time spent waiting for host budgets
"memo_hits", # answered from the response cache (max-age); nothing sent
"cache_hits", # scoreboard fetches answered from a shared ESPN cache entry
@@ -620,30 +616,6 @@ def _body_of(response: Any) -> Optional[bytes]:
return content if isinstance(content, bytes) else None
def _wire_bytes_of(response: Any, body: Optional[bytes]) -> int:
"""How many body bytes came off the socket for ``response``: the
compressed size when the server sent gzip, which ESPN does for every
scoreboard (63 KB on the wire for an 865 KB college football Saturday).
urllib3's ``HTTPResponse.tell()`` counts the raw bytes read before
decoding. A response without one (a test double, an adapter that is not
urllib3) or one whose body was not read is counted at its decoded size,
or as 0, so the counter never claims less than it can prove.
"""
if body is None:
return 0
raw = getattr(response, "raw", None)
tell = getattr(raw, "tell", None)
if callable(tell):
try:
read = tell()
except Exception:
read = None
if isinstance(read, int) and not isinstance(read, bool) and read > 0:
return read
return len(body)
def _retries_of(response: Any) -> int:
raw = getattr(response, "raw", None)
retries = getattr(raw, "retries", None)
@@ -1145,7 +1117,6 @@ class FetchService:
http_errors=int(status is not None and status >= 400),
retries=_retries_of(response),
bytes=len(body) if body is not None else 0,
wire_bytes=_wire_bytes_of(response, body),
throttled=int(waited > 0), overruns=int(overrun),
wait_seconds=waited)
except Exception:
+41
View File
@@ -135,6 +135,8 @@ import time
import traceback
from typing import Any, Callable, Dict, List, Optional, Tuple, TypedDict
from src.common import scroll_config
logger = logging.getLogger(__name__)
#: Bumped when a field changes meaning, so a reader can refuse stale files.
@@ -179,6 +181,12 @@ MAX_REFRESH_DROP = 0.2
#: trusted -- about a second of scrolling.
MIN_FRAMES_FOR_REFRESH = 90
#: Trusted windows, counting the one that adopted the period, before a panel
#: slower than its cap is reported. The estimate can still fall (the refresh
#: rate rise) by up to MAX_REFRESH_DROP per window early on; the warning
#: should not name a rate one more window would have corrected.
REFRESH_CHECK_WINDOWS = 3
FLUSH_INTERVAL = 10.0
#: A scroll's last frame older than this is a stall worth a stack dump.
@@ -441,6 +449,12 @@ class FrameTimingRecorder:
1.0 / refresh_hz if refresh_hz and refresh_hz > 0 else None)
# The first estimate, until a second window agrees with it.
self._refresh_candidate: Optional[float] = None
# The rate scroll speeds are solved against; see plan_refresh().
self.planned_refresh_hz: Optional[float] = None
# Trusted windows seen since the period was adopted, until the
# shortfall check has run.
self._refresh_windows = 0
self._shortfall_checked = True
self.totals: Dict[str, Any] = {
"static_frames": 0,
"scroll_frames": 0,
@@ -639,7 +653,12 @@ class FrameTimingRecorder:
self._refresh_candidate = estimate
elif current * (1.0 - MAX_REFRESH_DROP) <= estimate < current:
self.refresh_period = estimate
if self.refresh_period is not None:
self._refresh_windows += 1
period = self.refresh_period
if (period and not self._shortfall_checked
and self._refresh_windows >= REFRESH_CHECK_WINDOWS):
self._check_refresh_shortfall(1.0 / period)
histograms = self.histograms
for frame in batch:
@@ -688,6 +707,25 @@ class FrameTimingRecorder:
elif missed <= -1:
totals["early_frames"] += 1
def plan_refresh(self, hz: Optional[float]) -> None:
"""Say what rate scroll speeds are solved against, before frames arrive.
``DisplayManager.refresh_hz``: the configured cap. Once the measured
rate has held for :data:`REFRESH_CHECK_WINDOWS` windows, a panel that
falls short of it is logged once, with a cap it can hold (see
:func:`src.common.scroll_config.refresh_shortfall`). The display
manager calls this only for a real panel.
"""
self.planned_refresh_hz = hz
self._shortfall_checked = not hz
def _check_refresh_shortfall(self, measured_hz: float) -> None:
"""Log, once, a panel that cannot reach the rate speeds assume."""
self._shortfall_checked = True
shortfall = scroll_config.refresh_shortfall(measured_hz, self.planned_refresh_hz)
if shortfall:
logger.warning(scroll_config.describe_refresh_shortfall(shortfall))
def snapshot(self) -> Dict[str, Any]:
"""The JSON document: cumulative since this process started."""
if not self._binding_checked:
@@ -704,6 +742,9 @@ class FrameTimingRecorder:
"bucket_ms": BUCKET_MS,
"freeze_seconds": FREEZE_SECONDS,
"measured_refresh_hz": round(1.0 / period, 2) if period else None,
# Additive: what scroll speeds were solved against, so a reader can
# tell a stale file (written under another cap) from this one.
"planned_refresh_hz": self.planned_refresh_hz,
"binding_releases_gil": self._binding_gil,
"info": info,
"totals": copy.deepcopy(self.totals),
+63
View File
@@ -538,3 +538,66 @@ def speed_advice(
"smooth": smooth,
"alternatives": [as_dict(c) for c in alternatives],
}
#: A measured refresh this far below the rate speeds are planned for means
#: the panel cannot reach its cap. Smaller gaps are the cap's own slack and
#: the estimate's: one rig measured 99.95 Hz under a 100 Hz cap.
REFRESH_SHORTFALL = 0.03
#: How far under the measured rate a suggested cap sits. The measurement is
#: the fast end of the panel's refreshes (frame_timing takes the 10th
#: percentile of intervals), and an uncapped panel drifts: one read
#: 107.6-113.1 Hz over 15 seconds. A cap inside that band would not hold.
CAP_HEADROOM = 0.05
def holdable_cap(measured_hz: Any) -> Optional[int]:
"""A refresh cap the panel can hold: a multiple of 10, 5% under what it measured.
A multiple of 10 because its whole-pixel speeds are round numbers (a
100 Hz cap gives 50 and 100 px/s). None without a usable measurement, or
when the panel is too slow for any cap of 10 Hz or more.
"""
hz = _coerce(measured_hz)
if hz is None:
return None
cap = int(hz * (1.0 - CAP_HEADROOM) // 10) * 10
return cap if cap >= 10 else None
def refresh_shortfall(measured_hz: Any, planned_hz: Any) -> Optional[Dict[str, Any]]:
"""When the panel refreshes measurably slower than speeds are planned for.
``planned_hz`` is what :func:`configure` solves against -- the
``limit_refresh_rate_hz`` cap, or :data:`DEFAULT_REFRESH_HZ` when it is 0.
A panel that cannot reach it still moves whole pixels per frame, but every
speed runs slow by the shortfall and the ladder of smooth speeds is the
cap's, not the panel's. None when there is no measurement, or the panel
reaches the cap (or beats it, as some do by a few Hz).
"""
measured, planned = _coerce(measured_hz), _coerce(planned_hz)
if measured is None or planned is None:
return None
if measured >= planned * (1.0 - REFRESH_SHORTFALL):
return None
return {
"measured_hz": round(measured, 1),
"planned_hz": round(planned, 1),
"suggested_cap_hz": holdable_cap(measured),
"slow_percent": round((1.0 - measured / planned) * 100),
}
def describe_refresh_shortfall(shortfall: Dict[str, Any]) -> str:
"""One log line for :func:`refresh_shortfall`'s answer."""
text = (
f"The panel refreshes at about {shortfall['measured_hz']:.0f} Hz, below "
f"the {shortfall['planned_hz']:.0f} Hz that scroll speeds are planned "
f"for (display.hardware.limit_refresh_rate_hz), so every scroll runs "
f"about {shortfall['slow_percent']}% slower than configured and the "
f"smooth speeds are worked out for a rate this panel never reaches.")
if shortfall.get("suggested_cap_hz"):
text += (f" Set Limit Refresh Rate to {shortfall['suggested_cap_hz']} Hz "
f"(web UI, Display tab), which this panel can hold, and restart.")
return text
+39 -404
View File
@@ -1,56 +1,41 @@
"""What the panel shows next: the Arbiter of docs/RUN_LOOP_REDESIGN.md.
``Arbiter.decide(state, inputs, now)`` takes a snapshot that
``DisplayController.run()`` gathers and returns a :class:`ScreenPlan` naming
the Source that gets the panel. It is a pure function: no I/O, no clock
reads (``now`` is passed in), no locks, and it changes nothing it is given.
That is what lets a plain table of cases test the priority order, which
used to exist only as the order of ``if`` blocks in ``run()``.
``DisplayController.run()`` gathers once per pass and returns a
:class:`ScreenPlan` naming the Source that gets the panel. It is a pure
function: no I/O, no clock reads (``now`` is passed in), no locks, and it
changes nothing it is given. That is what lets a plain table of cases test
the priority order, which used to exist only as the order of ``if`` blocks
in ``run()``.
The order is
The full order is
ScheduledOff (a gate), Follower, OnDemand, Wifi, Live, Vegas, Rotation
Every Source but Vegas is decided here (stage 3). Vegas is the ``LEGACY``
plan: the Arbiter picks it, but its iteration is still run()'s own code
until stage 4.
Stage 2 decides the gate, Follower and Wifi. Every other case returns a
``LEGACY`` plan, meaning "carry on with run()'s existing code" (live
priority, Vegas, then one rotation screen). OnDemand is in the order already
because it outranks the WiFi notice: an active session is a ``LEGACY`` plan
even when a notice is pending.
A pass asks twice: once with the inputs every pass reads (the gate,
Follower, OnDemand, Wifi), and once more, only when nothing above the
notice took the panel, with the inputs the Sources below it need (whether
Vegas is on, the live-priority scan), read where run() always read them.
``decide(..., running=plan)`` is the other question, asked by the
ScreenRunner (src/screen_runner.py) at its service points: does a Source
in ``plan.preemptible_by`` now take the panel from the screen that is
running? The mid-screen rules are :func:`_hold_or_preempt`.
The state transitions (the next on-demand mode, a live claim and its
release, the rotation's step after a screen) are pure methods of
:class:`ArbiterState`; the controller applies what they return.
The Wifi Source's mid-screen rule, :func:`wifi_notice_preempts`, lives here
too, so both of its answers -- at the top of a pass and between frames --
come from one module.
"""
from dataclasses import dataclass, replace
from dataclasses import dataclass
from enum import Enum
from typing import FrozenSet, Optional, Protocol, Tuple
from typing import Optional
__all__ = [
"Arbiter",
"ArbiterInputs",
"ArbiterState",
"FramePolicy",
"LIVE_PREEMPTERS",
"SCHEDULED_OFF_DWELL",
"SCREEN_PREEMPTERS",
"ScreenEnd",
"ScreenPlan",
"Source",
"WIFI_NOTICE_DWELL",
"WifiNotice",
"live_pick",
"live_takeover",
"on_demand_bound",
"rotation_plan",
"wifi_notice_preempts",
]
@@ -68,41 +53,11 @@ class Source(Enum):
SCHEDULED_OFF = "scheduled-off"
FOLLOWER = "follower"
ON_DEMAND = "on-demand"
WIFI = "wifi"
LIVE = "live"
# Vegas: the Arbiter picks it, but its iteration is still run()'s own
# code (and its interrupt callback a second copy of this order) until
# stage 4 makes it a Source driven frame by frame.
# Not decided by the Arbiter yet: on-demand, live priority, Vegas and the
# rotation are still chosen by run()'s own code. Stage 3 adds the
# OnDemand, Live and Rotation Sources; stage 4 adds Vegas.
LEGACY = "legacy"
ROTATION = "rotation"
# Not a screen: a plugin reload waits at the top of the loop. It ends a
# screen between frames (the screen counts as shown and the rotation
# moves on), and the next pass reloads before it draws.
RELOAD = "reload"
class FramePolicy(Enum):
"""How often a screen draws: today's two frame loops (see
DisplayController._needs_high_fps). Stage 5 lets plugins declare it."""
#: The 125 Hz loop, paced to an 8 ms deadline: scrolling plugins.
HIGH_FPS = "high-fps"
#: The 1 Hz loop.
STATIC = "static"
class ScreenEnd(Protocol):
"""What ArbiterState.after needs to know about how a screen ended
(screen_runner.Outcome, filled in by the controller)."""
@property
def on_demand_active(self) -> bool:
"""An on-demand session was running when the screen ended."""
@property
def still_live(self) -> bool:
"""The mode's plugin still had live content: hold the rotation."""
@dataclass(frozen=True)
@@ -121,122 +76,11 @@ class WifiNotice:
class ArbiterState:
"""What the Arbiter remembers between passes.
A snapshot of the controller's own fields, taken when decide() is
called (DisplayController._arbiter_state); the transitions below return
the next state, and the controller writes it back.
Attributes:
current_mode: The mode on the panel or about to be
(``current_display_mode``).
on_demand_modes: The on-demand session's modes, in the order it
shows them (a pinned mode already moved to the front when the
session started, by _apply_on_demand_pin).
on_demand_index: Which of them is showing.
on_demand_expires_at: When the session ends (wall clock), or None
for a session with no duration.
on_demand_pinned: The session was started pinned. Carried for the
snapshot; the pin itself is already in ``on_demand_modes``.
rotation: The rotation's modes (``available_modes``).
rotation_index: Where the rotation is (``current_mode_index``).
live_resume_index: Where the rotation was when live priority took
the panel, so it resumes there once nothing is live; None while
live priority holds nothing.
live_takeover_unshown: A mid-screen takeover chose current_mode and
it has not been shown yet, so the next pass must not advance the
live round-robin past it.
Nothing yet: the stage-2 Sources decide from the inputs alone. The
on-demand index, the rotation index and the live resume point move here
with their Sources in stage 3.
"""
current_mode: Optional[str] = None
on_demand_modes: Tuple[str, ...] = ()
on_demand_index: int = 0
on_demand_expires_at: Optional[float] = None
on_demand_pinned: bool = False
rotation: Tuple[str, ...] = ()
rotation_index: int = 0
live_resume_index: Optional[int] = None
live_takeover_unshown: bool = False
def next_on_demand(self) -> "ArbiterState":
"""The session's next mode, wrapping round. Needs a mode list."""
index = (self.on_demand_index + 1) % len(self.on_demand_modes)
return replace(self, on_demand_index=index,
current_mode=self.on_demand_modes[index])
def claim_live(self, mode: str) -> "ArbiterState":
"""Live priority takes the panel for ``mode``.
The rotation's position is saved only on the first claim, not on
each re-check while the hold continues, so it resumes where live
priority interrupted it instead of after the live mode (which would
skip every mode between the two).
"""
if self.current_mode == mode:
return self
resume = self.rotation_index if self.live_resume_index is None else self.live_resume_index
index = self.rotation.index(mode) if mode in self.rotation else self.rotation_index
return replace(self, current_mode=mode, rotation_index=index,
live_resume_index=resume)
def after(self, outcome: "ScreenEnd") -> "ArbiterState":
"""The state once a screen has run its course: the next mode.
An on-demand session moves to its next mode. Otherwise the rotation
advances -- unless the mode just shown is a live-priority mode that
is still live, which holds the panel. A session with no modes left
is ended by the controller before it asks (that is not pure: it
resumes the rotation and clears the cache).
"""
if outcome.on_demand_active:
return self.next_on_demand() if self.on_demand_modes else self
if outcome.still_live or not self.rotation:
return self
index = (self.rotation_index + 1) % len(self.rotation)
return replace(self, rotation_index=index, current_mode=self.rotation[index])
def release_live(self) -> "ArbiterState":
"""Nothing is live any more: the rotation resumes where it was."""
if self.live_resume_index is None or not self.rotation:
return self
index = self.live_resume_index % len(self.rotation)
return replace(self, current_mode=self.rotation[index], rotation_index=index,
live_resume_index=None)
def showing(self, plan: "ScreenPlan") -> "ArbiterState":
"""The state once ``plan`` is on the panel.
An on-demand plan puts the session's index on the mode it shows (an
index past the end of a shortened list starts it again at 0).
"""
state = replace(self, current_mode=plan.mode)
if plan.source is Source.ON_DEMAND and self.on_demand_modes:
state = replace(state, on_demand_index=_on_demand_index(self))
return state
def _on_demand_index(state: ArbiterState) -> int:
"""The session's index, or 0 once it is past the end of its list."""
index = state.on_demand_index
return index if index < len(state.on_demand_modes) else 0
def on_demand_bound(min_duration: float, max_duration: float,
deadline: Optional[float],
now: float) -> Optional[Tuple[float, float]]:
"""Shorten a screen's (min, max) seconds to what is left of a timed
on-demand session ending at ``deadline``. None when nothing is left.
The OnDemand Source's bound, applied after the screen's first frame,
where it always was (``now`` is read then).
"""
if deadline is None:
return min_duration, max_duration
remaining = max(0.0, deadline - now)
min_duration = min(min_duration, remaining)
max_duration = min(max_duration, remaining)
if max_duration <= 0:
return None
return min_duration, max_duration
@dataclass(frozen=True)
class ArbiterInputs:
@@ -251,122 +95,57 @@ class ArbiterInputs:
when it could win (the panel is on, and neither a follower nor
on-demand outranks it), because reading it has side effects: a
1 Hz throttle and deleting an expired file.
live_modes: The modes with live content, from a live-priority scan,
in registration order; None when no scan was made (on-demand,
Vegas keeping live content in its ticker, a throttled
mid-screen check). A scan asks every live-priority plugin, so it
is made only where run() always made it.
vegas_enabled: Vegas mode is on (and no on-demand session holds it
off).
vegas_live_in_ticker: Vegas keeps live content in its ticker
instead of yielding the panel to it.
vegas_yielded: This pass's Vegas iteration has run and yielded, so
the Vegas Source passes and the screen it fell through to is
decided.
reload_pending: Mid-screen only: a plugin reload is waiting for the
top of the loop, at a service point where that ends the screen.
"""
schedule_on: bool
on_demand_active: bool
follower_active: bool
wifi_notice: Optional[WifiNotice] = None
live_modes: Optional[Tuple[str, ...]] = None
vegas_enabled: bool = False
vegas_live_in_ticker: bool = False
vegas_yielded: bool = False
reload_pending: bool = False
@dataclass(frozen=True)
class ScreenPlan:
"""The Arbiter's answer for one pass.
decide() is pure, so it cannot ask a plugin anything: the fields a
plugin answers (its durations, whether it runs a dynamic cycle, how
often it draws) are filled in by the controller after the screen's first
frame, when they have always been read (DisplayController.complete_plan).
Attributes:
source: The Source that gets the panel.
mode: The display mode to draw (None for a blank, follower or notice).
plugin: The id of the plugin drawing ``mode``, once resolved.
min_duration: Seconds the screen runs at least (dynamic duration).
max_duration: How long the plan holds the panel, in seconds, at most
(its dwell ends early when what the panel should show changes).
None when the Source paces itself (a follower frame, Vegas), and
for a rotation or live plan until its first frame.
dynamic: Run until the plugin's cycle completes, between min and max.
frame_policy: Which frame loop the screen runs.
preemptible_by: The Sources that may end the screen mid-way.
None when the Source paces itself: a follower frame, or LEGACY.
notice: The WiFi notice to draw, for a WIFI plan.
deadline: For an on-demand plan, when the session ends (wall
clock): after the first frame the screen's durations are cut to
what is left (:func:`on_demand_bound`).
ends_live: Nothing is live any more and live priority had
interrupted the rotation: taking this plan resumes the rotation
where it was (ArbiterState.release_live) before it shows.
"""
source: Source
mode: Optional[str] = None
plugin: Optional[str] = None
min_duration: Optional[float] = None
max_duration: Optional[float] = None
dynamic: bool = False
frame_policy: Optional[FramePolicy] = None
preemptible_by: FrozenSet[Source] = frozenset()
notice: Optional[WifiNotice] = None
deadline: Optional[float] = None
ends_live: bool = False
SCHEDULED_OFF_PLAN = ScreenPlan(Source.SCHEDULED_OFF, max_duration=SCHEDULED_OFF_DWELL)
FOLLOWER_PLAN = ScreenPlan(Source.FOLLOWER)
RELOAD_PLAN = ScreenPlan(Source.RELOAD)
#: What may end a screen mid-way: the schedule, an on-demand session
#: starting or ending, a WiFi notice, a live game, the rotation moving
#: under the screen, and a plugin reload. Not a follower or Vegas: those
#: are only looked at between screens.
SCREEN_PREEMPTERS: FrozenSet[Source] = frozenset(
{Source.SCHEDULED_OFF, Source.ON_DEMAND, Source.WIFI, Source.LIVE, Source.ROTATION,
Source.RELOAD})
#: A live screen is not preempted by Live: live games take turns between
#: screens, never mid-screen.
LIVE_PREEMPTERS: FrozenSet[Source] = SCREEN_PREEMPTERS - {Source.LIVE}
LEGACY_PLAN = ScreenPlan(Source.LEGACY)
class Arbiter:
"""Decides which Source gets the panel. Stateless; see the module docstring."""
@staticmethod
def decide(state: ArbiterState, inputs: ArbiterInputs, now: float,
running: Optional[ScreenPlan] = None) -> ScreenPlan:
def decide(state: ArbiterState, inputs: ArbiterInputs, now: float) -> ScreenPlan:
"""The plan for this pass, from the Sources in priority order.
With ``running``, the question is the ScreenRunner's at one of its
service points instead: does a Source in ``running.preemptible_by``
now take the panel from that screen? The answer is ``running``
itself (the same object) while it holds, else the plan that ends it.
See :func:`_hold_or_preempt` for the rules.
Args:
state: What the Arbiter remembers between passes.
state: What the Arbiter remembers between passes (nothing yet).
inputs: This pass's snapshot.
now: Wall-clock time of the snapshot. The OnDemand Source reads
it for what is left of a timed session, and the mid-screen
WiFi rule (:func:`wifi_notice_preempts`) to compare with the
notice's expiry. The top-of-pass WiFi check does not: it
takes the notice as read.
running: The screen on the panel, for a mid-screen check.
now: Wall-clock time of the snapshot. No stage-2 Source reads it:
the top-of-pass WiFi check takes the notice as read, and only
the mid-screen check (:func:`wifi_notice_preempts`) compares
it with the expiry. It is in the signature for the Sources
stage 3 adds (on-demand expiry, durations).
Returns:
The winning Source's plan (LEGACY for Vegas), or ``running``.
The winning Source's plan, or LEGACY_PLAN when the winner is one
run() still decides itself.
"""
if running is not None:
return _hold_or_preempt(state, inputs, now, running)
del state, now # not read by the stage-2 Sources; see the docstring
# ScheduledOff is a gate, not a Source: a scheduled-off panel stays
# blank even for a follower, and only an on-demand session overrides
@@ -378,161 +157,17 @@ class Arbiter:
if inputs.follower_active:
return FOLLOWER_PLAN
# 2. OnDemand: the session's current mode. It outranks the notice.
# 2. OnDemand: decided by run() until stage 3. It outranks the notice.
if inputs.on_demand_active:
return _on_demand_plan(state, now)
return LEGACY_PLAN
# 3. Wifi: a pending notice, held for one short dwell per pass.
if inputs.wifi_notice is not None:
return ScreenPlan(Source.WIFI, max_duration=WIFI_NOTICE_DWELL,
notice=inputs.wifi_notice)
# 4. Live: the next live game, round-robin across several. With
# nothing live, a rotation that live priority interrupted resumes.
ends_live = False
if _live_applies(inputs):
pick = live_pick(inputs.live_modes, state.current_mode,
advance=not state.live_takeover_unshown)
if pick is not None:
return ScreenPlan(Source.LIVE, mode=pick, preemptible_by=LIVE_PREEMPTERS)
ends_live = state.live_resume_index is not None and bool(state.rotation)
# 5. Vegas: one iteration of the ticker, run by run()'s own code
# until stage 4. Passes once this pass's iteration has yielded.
if inputs.vegas_enabled and not inputs.vegas_yielded:
return ScreenPlan(Source.LEGACY, ends_live=ends_live)
# 6. Rotation: the rotation's current mode (after the resume, when
# live priority just ended).
return rotation_plan(state.release_live() if ends_live else state,
ends_live=ends_live)
def _hold_or_preempt(state: ArbiterState, inputs: ArbiterInputs, now: float,
running: ScreenPlan) -> ScreenPlan:
"""The mid-screen rules: ``running``, or the plan that ends it.
What the frame loops used to check one by one (_check_live_takeover,
then _screen_preempted with _wifi_notice_pending in it, before stage 3),
in their order:
1. Live: a game went live while a non-live screen runs (the inputs
carry a scan only when one was due, at most once a second). Checked
first because it is the one preemption that changes the state -- the
rotation moves to the live mode and remembers where it was -- and it
still happens when a WiFi notice is also pending: the next pass then
shows the notice, and the game after it.
2. The panel's mode moved under the screen: an on-demand session
started, ended or changed mode, or the rotation was rebuilt (a
plugin enabled, disabled or reloaded).
3. The schedule turned the panel off.
4. A WiFi notice arrived (unless on-demand outranks it), compared with
its expiry because the read throttle can hand back a stale one.
5. A plugin reload is waiting at the top of the loop.
A follower and Vegas are never mid-screen preemptions; they are looked
at between screens.
"""
by = running.preemptible_by
if Source.LIVE in by:
takeover = live_takeover(state, inputs)
if takeover is not None:
return ScreenPlan(Source.LIVE, mode=takeover, preemptible_by=LIVE_PREEMPTERS)
if state.current_mode != running.mode:
source = Source.ON_DEMAND if inputs.on_demand_active else Source.ROTATION
if source in by:
return ScreenPlan(source, mode=state.current_mode, preemptible_by=SCREEN_PREEMPTERS)
if (Source.SCHEDULED_OFF in by and not inputs.schedule_on
and not inputs.on_demand_active):
return SCHEDULED_OFF_PLAN
notice = inputs.wifi_notice
if (Source.WIFI in by and notice is not None
and wifi_notice_preempts(notice, inputs.on_demand_active, now)):
return ScreenPlan(Source.WIFI, max_duration=WIFI_NOTICE_DWELL, notice=notice)
if Source.RELOAD in by and inputs.reload_pending:
return RELOAD_PLAN
return running
def live_takeover(state: ArbiterState, inputs: ArbiterInputs) -> Optional[str]:
"""The live mode that takes the panel mid-screen, or None.
The first live mode, when a scan found one and the panel is not on a
live mode already. Never while on-demand holds the panel, while it is
scheduled off, or while Vegas keeps live content in its ticker.
"""
if not _live_applies(inputs) or inputs.on_demand_active or not inputs.schedule_on:
return None
live = inputs.live_modes
if not live or state.current_mode in live:
return None
return live[0]
def _on_demand_plan(state: ArbiterState, now: float) -> ScreenPlan:
"""The OnDemand Source: the session's current mode.
``max_duration`` is what is left of a timed session at ``now`` (None
without a duration); ``deadline`` carries the expiry so the bound can be
applied again after the first frame. A session with no modes left (its
plugin was unloaded under it) gets a plan with no mode: the controller
ends the session and shows the rotation's mode instead.
"""
modes = state.on_demand_modes
if not modes:
return ScreenPlan(Source.ON_DEMAND)
expires_at = state.on_demand_expires_at
remaining = None if expires_at is None else max(0.0, expires_at - now)
return ScreenPlan(Source.ON_DEMAND, mode=modes[_on_demand_index(state)],
max_duration=remaining, deadline=expires_at,
preemptible_by=SCREEN_PREEMPTERS)
def rotation_plan(state: ArbiterState, ends_live: bool = False) -> ScreenPlan:
"""The Rotation Source: the mode the rotation is on.
That is ``state.current_mode``, which is ``rotation[rotation_index]``
except where something moved the panel off the list and the rotation
carries on from there: a live mode no rotation entry names, or None
when a session ended with no enabled mode to resume to.
"""
return ScreenPlan(Source.ROTATION, mode=state.current_mode, ends_live=ends_live,
preemptible_by=SCREEN_PREEMPTERS)
def _live_applies(inputs: ArbiterInputs) -> bool:
"""Whether the Live Source has a say: a scan was made, and Vegas is not
keeping live content in its ticker (where the live plugin takes extra
turns in the marquee instead of the panel)."""
if inputs.live_modes is None:
return False
return not (inputs.vegas_enabled and inputs.vegas_live_in_ticker)
def live_pick(live_modes: Optional[Tuple[str, ...]], current_mode: Optional[str],
advance: bool) -> Optional[str]:
"""The live mode to show, or None when nothing is live.
When several plugins are live at once this round-robins between them, so
the panel alternates each dwell instead of pinning to the first one
registered. The mode on the panel is the cursor, so this stays right as
games start and end.
Args:
live_modes: The live modes, in registration order.
current_mode: The mode on the panel.
advance: True for the rotation's pick (the live mode after the one
showing). False for a peek (the one showing if it is still live,
else the first), which Vegas uses to ask whether anything is.
"""
if not live_modes:
return None
if current_mode in live_modes:
if advance:
index = live_modes.index(current_mode)
return live_modes[(index + 1) % len(live_modes)]
return current_mode
return live_modes[0]
# 4-6. Live, Vegas, Rotation: still run()'s own code.
return LEGACY_PLAN
def wifi_notice_preempts(notice: Optional[WifiNotice], on_demand_active: bool,
+526 -665
View File
File diff suppressed because it is too large Load Diff
+8 -1
View File
@@ -393,6 +393,11 @@ class DisplayManager:
self._setup_matrix()
logger.info("Matrix setup completed in %.3f seconds", time.time() - start_time)
# Only a real panel's swaps wait on its refresh: the emulator and the
# fallback canvas pace themselves, so "slower than the cap" would be
# noise there.
if self.matrix is not None and os.environ.get('EMULATOR', 'false') != 'true':
self.frame_timing.plan_refresh(self.refresh_hz)
self._setup_scan_order_compensation()
font_time = time.time()
@@ -1501,7 +1506,9 @@ class DisplayManager:
fractional-pixel motion. See src/common/scroll_config.py.
Note this is the configured *cap*, not necessarily what the panel
achieves -- scripts/scroll_speeds.py --measure reports the real rate.
achieves -- scripts/scroll_speeds.py --measure reports the real rate,
and the frame-timing recorder logs a warning, with a cap the panel can
hold, once it has measured a panel that falls short of this.
"""
hardware = (self.config.get('display') or {}).get('hardware') or {}
try:
+31 -125
View File
@@ -490,13 +490,10 @@ def record_error(
# user reads, and the other way round for the clear request.
#
# ERROR_SNAPSHOT_KEY written by the display service only
# ERROR_CLEAR_REQUEST_KEY written by the web interface only, as a fallback
# ERROR_CLEAR_REQUEST_KEY written by the web interface only
#
# A clear goes over the control socket (``errors.clear``): the display applies
# it (clear_before) and republishes the snapshot before it answers. Only when
# the socket cannot carry it (no socket, or a display older than the command)
# does the web interface record a request in the mailbox, which the display
# applies on its next tick; its tick reads that file only when it changed.
# A clear is asynchronous: the web interface records a request, and the
# display service applies it (clear_before) on its next tick and republishes.
# Until it has, the web interface hides whatever the snapshot shows from
# before the cutoff, so a clear takes effect for readers immediately and a
# snapshot published just before the request cannot bring old errors back.
@@ -601,43 +598,13 @@ class ErrorSnapshotPublisher:
self._published_version: Optional[int] = None
self._last_attempt: Optional[float] = None
self._applied_clear_id: Optional[str] = None
# The widest cutoff applied in this process: a clear request at or
# before it has nothing left to clear (see _pending_cutoff).
self._applied_clear_cutoff: Optional[float] = None
from src.cache_manager import MailboxWatch # the display's cache, loaded already
self._mailbox = MailboxWatch(ERROR_CLEAR_REQUEST_KEY)
self._tick_lock = threading.Lock()
self._stop = threading.Event()
self._thread: Optional[threading.Thread] = None
def _clear(self, request_id: str, cutoff: float) -> int:
"""Apply one clear and remember it. Caller holds _tick_lock."""
cleared = 0
if math.isfinite(cutoff):
cleared = self.aggregator.clear_before(datetime.fromtimestamp(cutoff))
_snapshot_logger.info("Cleared %d plugin error record(s) as requested (%s)",
cleared, request_id)
if self._applied_clear_cutoff is None or cutoff > self._applied_clear_cutoff:
self._applied_clear_cutoff = cutoff
# A malformed request is acknowledged too, so it is not retried forever.
self._applied_clear_id = request_id
return cleared
def _apply_clear_request(self) -> bool:
"""Honour a mailbox clear request we have not applied yet. True if one was.
The mailbox is the fallback for a web interface that could not use
the control socket (``errors.clear``, :meth:`clear_now`). It is read
only when its file changed since the last tick; otherwise a tick
costs one stat().
"""
if not self._mailbox.changed(self.cache_manager):
return False
try:
request = self.cache_manager.get(ERROR_CLEAR_REQUEST_KEY, max_age=None, memory_ttl=0)
except Exception:
self._mailbox.forget()
raise
"""Honour a clear request we have not applied yet. True if one was."""
request = self.cache_manager.get(ERROR_CLEAR_REQUEST_KEY, max_age=None, memory_ttl=0)
if not isinstance(request, dict):
return False
request_id = request.get("request_id")
@@ -647,30 +614,14 @@ class ErrorSnapshotPublisher:
cutoff = float(request.get("cutoff"))
except (TypeError, ValueError):
cutoff = float("nan")
self._clear(request_id, cutoff)
if math.isfinite(cutoff):
cleared = self.aggregator.clear_before(datetime.fromtimestamp(cutoff))
_snapshot_logger.info("Cleared %d plugin error record(s) as requested (%s)",
cleared, request_id)
# A malformed request is acknowledged too, so it is not retried forever.
self._applied_clear_id = request_id
return True
def clear_now(self, request_id: str, cutoff: float) -> int:
"""``errors.clear`` over the control socket: apply a clear at once and
republish the snapshot, so the web interface's next read has it.
Returns how many records were cleared. Raises when the snapshot
could not be written, so the caller is not told it worked."""
with self._tick_lock:
cleared = self._clear(request_id, float(cutoff))
self._publish(self.aggregator.version, self._clock())
return cleared
def _publish(self, version: int, now: float) -> None:
"""Write the snapshot. Caller holds _tick_lock."""
# Stamp the attempt before writing: a cache that keeps failing
# is retried at the throttled rate, not on every tick.
self._last_attempt = now
snapshot = self.aggregator.build_snapshot()
snapshot["applied_clear_id"] = self._applied_clear_id
snapshot["applied_clear_cutoff"] = self._applied_clear_cutoff
self.cache_manager.set(ERROR_SNAPSHOT_KEY, snapshot)
self._published_version = version
def tick(self) -> bool:
"""Apply a pending clear and publish if due. True if a snapshot was written."""
with self._tick_lock:
@@ -684,7 +635,13 @@ class ErrorSnapshotPublisher:
if (self._last_attempt is not None
and now - self._last_attempt < self.min_interval):
return False
self._publish(version, now)
# Stamp the attempt before writing: a cache that keeps failing
# is retried at the throttled rate, not on every tick.
self._last_attempt = now
snapshot = self.aggregator.build_snapshot()
snapshot["applied_clear_id"] = self._applied_clear_id
self.cache_manager.set(ERROR_SNAPSHOT_KEY, snapshot)
self._published_version = version
return True
except Exception as err: # never let reporting break the display
_snapshot_logger.debug("Could not publish the plugin error snapshot: %s",
@@ -736,22 +693,6 @@ def start_error_snapshot_publisher(cache_manager: Any) -> Optional[ErrorSnapshot
return None
def apply_error_clear(request_id: str, args: Any) -> Dict[str, Any]:
"""The display's handler for ``errors.clear`` on the control socket.
``args`` is the contract's ErrorsClearArgs (``cutoff``, epoch seconds).
Runs on the socket's connection thread: the aggregator and the publisher
have their own locks, and nothing here touches rendering. Returns
ErrorsClearResult once the clear is applied and the snapshot rewritten.
"""
publisher = _snapshot_publisher
if publisher is None:
raise RuntimeError("the error snapshot publisher is not running")
cutoff = float(args.cutoff)
cleared = publisher.clear_now(request_id, cutoff)
return {"request_id": request_id, "cutoff": cutoff, "cleared": cleared}
# --- Reading side (web interface) -------------------------------------------
def read_error_report(cache_manager: Any) -> Tuple[Optional[Dict[str, Any]], Optional[Dict[str, Any]]]:
@@ -790,15 +731,7 @@ def _pending_cutoff(snapshot: Optional[Dict[str, Any]],
cutoff = float(clear_request.get("cutoff"))
except (TypeError, ValueError):
return None
if not math.isfinite(cutoff):
return None
# A wider clear has been applied since (over the control socket): this
# older request has nothing left to hide.
applied = snapshot.get("applied_clear_cutoff") if snapshot is not None else None
if (isinstance(applied, (int, float)) and not isinstance(applied, bool)
and applied >= cutoff):
return None
return cutoff
return cutoff if math.isfinite(cutoff) else None
def _is_after(item: Any, field_name: str, cutoff: float) -> bool:
@@ -898,31 +831,13 @@ def _count_cleared(summary: Dict[str, Any], cutoff: float) -> Optional[int]:
return None
#: ``send(request_id, cutoff)`` hands a clear to the display over the control
#: socket and returns its ErrorsClearResult, or None when the socket could
#: not carry it and the mailbox should be written instead. Any exception it
#: raises reaches the caller: the display had the request and failed it.
ClearSender = Callable[[str, float], Optional[Dict[str, Any]]]
def request_error_clear(cache_manager: Any, cutoff: float,
send: Optional[ClearSender] = None) -> Dict[str, Any]:
def request_error_clear(cache_manager: Any, cutoff: float) -> Dict[str, Any]:
"""Ask the display service to forget errors recorded at or before ``cutoff``.
Over the control socket when ``send`` is given and carries it: the
display applies the clear and republishes its snapshot before it
answers, so nothing is written here. Otherwise (no socket, or a display
older than ``errors.clear``) a request is written to the
``plugin_error_clear_request`` mailbox, which the display applies on
its next tick, and readers hide the cleared errors until then.
Returns ``request_id``, ``cutoff`` (ISO, local time), ``cleared_count``,
``clear_requested``, ``applied`` (the display has already cleared them)
and ``transport`` (``socket`` or ``mailbox``). ``cleared_count`` is the
display's own count over the socket, else an estimate from the snapshot
(see _count_cleared). Raises OSError when a mailbox request did not reach
the shared cache, since a cache without a usable directory accepts set()
and keeps nothing.
Returns ``request_id``, ``cutoff`` (ISO, local time), ``cleared_count``
(see _count_cleared) and ``clear_requested``. Raises OSError when the
request did not reach the shared cache, since a cache without a usable
directory accepts set() and keeps nothing.
A request the display has not applied yet is only ever widened: a later,
narrower one ("older than 24 hours" after "everything") overwriting it
@@ -933,21 +848,8 @@ def request_error_clear(cache_manager: Any, cutoff: float,
if pending is not None:
cutoff = max(cutoff, pending)
before = error_summary_from_report(snapshot, clear_request)
request_id = uuid.uuid4().hex
answer = {
"clear_requested": True,
"request_id": request_id,
"cutoff": datetime.fromtimestamp(cutoff).isoformat(),
}
if send is not None:
result = send(request_id, cutoff)
if result is not None:
count = result.get("cleared")
return dict(answer, applied=True, transport="socket",
cleared_count=count if isinstance(count, int) and not isinstance(count, bool)
else _count_cleared(before, cutoff))
request = {
"request_id": request_id,
"request_id": uuid.uuid4().hex,
"cutoff": cutoff,
"requested_at": time.time(),
}
@@ -955,5 +857,9 @@ def request_error_clear(cache_manager: Any, cutoff: float,
stored = cache_manager.get(ERROR_CLEAR_REQUEST_KEY, max_age=None, memory_ttl=0)
if not isinstance(stored, dict) or stored.get("request_id") != request["request_id"]:
raise OSError("the clear request was not stored in the shared cache")
return dict(answer, applied=False, transport="mailbox",
cleared_count=_count_cleared(before, cutoff))
return {
"cleared_count": _count_cleared(before, cutoff),
"clear_requested": True,
"request_id": request["request_id"],
"cutoff": datetime.fromtimestamp(cutoff).isoformat(),
}
+10 -71
View File
@@ -3,13 +3,8 @@
Every failure -- no socket (the display is stopped, or predates the socket),
a refused or timed-out connection, a reply that breaks the contract, or an
error the display returned -- raises :class:`ControlError` with a short
``reason``. Nothing here blocks for longer than ``timeout`` in total.
Whether the caller may then write the file mailbox instead is
:func:`should_fall_back`: only when the display never took the request (it
could not be reached, or it is too old to know the command). A display that
took the request and then failed, refused or went quiet is answered as
that, not posted a second time through the mailbox.
``reason``, and the caller falls back to the file mailbox. Nothing here
blocks for longer than ``timeout`` in total.
"""
from __future__ import annotations
@@ -27,7 +22,6 @@ from src.ipc.contract import (
SUBSCRIBE_KEEPALIVE_SECONDS,
SUPPORTED_VERSIONS,
Command,
ErrorCode,
FrameReader,
ProtocolError,
Request,
@@ -55,49 +49,17 @@ class ControlError(Exception):
``refused``, ``timeout``, ``closed``, ``bad_response``, ``invalid_request``.
When the display answered with an error, ``reason`` is that error's
:class:`~src.ipc.contract.ErrorCode` (``busy``, ``unknown_command``, ...).
``sent`` is True once the whole request was written to a connected
display, which may then have acted on it. A refusal the display sends
before it reads anything (``forbidden``, too many connections) carries
no request id and leaves ``sent`` False.
"""
def __init__(self, reason: str, message: str = '', *, sent: bool = False):
def __init__(self, reason: str, message: str = ''):
super().__init__(reason, message)
self.reason = reason
self.message = message
self.sent = sent
def __str__(self) -> str:
return f'{self.reason}: {self.message}' if self.message else self.reason
#: Answers from a display that read the request but does not speak it: one
#: older than the command (an upgrade in progress) or the protocol version.
#: It did nothing, so the mailbox is the way to reach it.
UPGRADE_REASONS = frozenset({ErrorCode.UNKNOWN_COMMAND, ErrorCode.UNSUPPORTED_VERSION})
def should_fall_back(error: BaseException) -> bool:
"""May the caller write the file mailbox after ``error``?
Yes when the display never took the request: there is no socket (the
display is stopped, predates the socket, or it is switched off), the
connection was refused or timed out, the display turned the connection
away before reading it, or it is too old to know the command
(:data:`UPGRADE_REASONS`). Also for an error that is not a
:class:`ControlError` (a bug in the client), as before.
No once the display had the request: a ``busy`` queue, ``invalid_args``,
an ``internal`` error, or a timeout or hang-up after the request was
sent. The display may have applied it, or would refuse it from the
mailbox too, so a second copy there only hides the failure.
"""
if not isinstance(error, ControlError):
return True
return not error.sent or error.reason in UPGRADE_REASONS
def request(cmd: str, args: Optional[Mapping[str, Any]] = None, *,
request_id: Optional[str] = None,
timeout: float = DEFAULT_TIMEOUT_SECONDS,
@@ -131,14 +93,11 @@ def request(cmd: str, args: Optional[Mapping[str, Any]] = None, *,
# A refusal before the request was read (forbidden, too many
# connections) carries no id.
if response.id != request_id and not (response.id is None and not response.ok):
raise ControlError('bad_response', 'the reply is for a different request', sent=True)
raise ControlError('bad_response', 'the reply is for a different request')
if not response.ok:
error = response.error
# No id: refused at the door (forbidden, too many connections),
# before the display read the request.
raise ControlError(error.code if error else 'bad_response',
error.message if error else '',
sent=response.id is not None)
error.message if error else '')
return dict(response.result or {})
@@ -183,31 +142,26 @@ def _connect(paths: Sequence[str], deadline: float) -> socket.socket:
def _exchange(sock: socket.socket, payload: bytes, deadline: float) -> Response:
"""Send ``payload`` and read the reply. A failure once the whole request
is written raises with ``sent=True``: the display may have it."""
sent = False
try:
sock.settimeout(_remaining(deadline))
sock.sendall(payload)
sent = True
reader = FrameReader(MAX_MESSAGE_BYTES)
while True:
sock.settimeout(_remaining(deadline))
data = sock.recv(4096)
if not data:
raise ControlError('closed', 'the display closed the connection', sent=sent)
raise ControlError('closed', 'the display closed the connection')
lines = reader.feed(data)
if lines:
return Response.from_dict(decode_message(lines[0]))
except socket.timeout:
raise ControlError('timeout', 'no reply in time', sent=sent) from None
raise ControlError('timeout', 'no reply in time') from None
except ProtocolError as e:
raise ControlError('bad_response', e.message, sent=sent) from None
except ControlError as e:
e.sent = e.sent or sent
raise ControlError('bad_response', e.message) from None
except ControlError:
raise
except OSError as e:
raise ControlError('closed', str(e), sent=sent) from None
raise ControlError('closed', str(e)) from None
# -- commands ---------------------------------------------------------------------------
@@ -273,21 +227,6 @@ def plugin_reload(plugin_id: str, *, timeout: Optional[float] = None,
else timeout, paths=paths)
def errors_clear(request_id: str, cutoff: float, *,
timeout: float = DEFAULT_TIMEOUT_SECONDS,
paths: Optional[Sequence[str]] = None) -> Dict[str, Any]:
"""Have the display forget the plugin errors recorded at or before
``cutoff`` (epoch seconds) and publish its error snapshot again.
Returns :class:`~src.ipc.contract.ErrorsClearResult` once it is done.
Raises :class:`ControlError`: ``unknown_command`` from a display older
than the command, which still reads the ``plugin_error_clear_request``
mailbox.
"""
return request(Command.ERRORS_CLEAR, {'cutoff': cutoff}, request_id=request_id,
timeout=timeout, paths=paths)
def ping(*, timeout: float = DEFAULT_TIMEOUT_SECONDS,
paths: Optional[Sequence[str]] = None) -> Dict[str, Any]:
return request(Command.PING, {}, timeout=timeout, paths=paths)
+4 -44
View File
@@ -149,13 +149,12 @@ class Command:
PLUGIN_RELOAD = 'plugin.reload'
STATE_GET = 'state.get'
STATE_SUBSCRIBE = 'state.subscribe'
ERRORS_CLEAR = 'errors.clear'
#: Every command version 1 defines, in the order ``hello`` reports them.
#: ``brightness.set`` and ``plugin.reload`` came in stage 2, ``state.get``
#: and ``state.subscribe`` in stage 3, and ``errors.clear`` in stage 4, all
#: within version 1 (see the module docstring on adding commands).
#: ``brightness.set`` and ``plugin.reload`` came in stage 2, and ``state.get``
#: and ``state.subscribe`` in stage 3, all within version 1 (see the module
#: docstring on adding commands).
COMMANDS: Tuple[str, ...] = (
Command.HELLO,
Command.PING,
@@ -166,15 +165,8 @@ COMMANDS: Tuple[str, ...] = (
Command.PLUGIN_RELOAD,
Command.STATE_GET,
Command.STATE_SUBSCRIBE,
Command.ERRORS_CLEAR,
)
#: Commands the connection thread answers itself, through a handler the
#: display registers (``ControlServer(handlers=...)``), because they touch
#: nothing the render thread owns. A display that registered none answers
#: ``unknown_command``, and the client falls back as from an older display.
DIRECT_COMMANDS = frozenset({Command.ERRORS_CLEAR})
#: Commands that are queued for the render thread.
QUEUED_COMMANDS = frozenset({Command.ON_DEMAND_START, Command.ON_DEMAND_STOP,
Command.BRIGHTNESS_SET, Command.PLUGIN_RELOAD})
@@ -573,32 +565,8 @@ class StateSubscribeArgs:
return cls()
@dataclass(frozen=True)
class ErrorsClearArgs:
"""``errors.clear``: forget the plugin errors recorded at or before
``cutoff`` (seconds since the epoch), as ``POST /api/v3/errors/clear``
asks. The request id is the clear's id, which the display's error
snapshot then reports as ``applied_clear_id``.
"""
cutoff: float
def to_dict(self) -> Dict[str, Any]:
return {'cutoff': self.cutoff}
@classmethod
def from_dict(cls, args: Mapping[str, Any]) -> 'ErrorsClearArgs':
value = args.get('cutoff')
if isinstance(value, bool) or not isinstance(value, (int, float)):
raise ProtocolError(ErrorCode.INVALID_ARGS, 'cutoff must be a number of seconds')
if not math.isfinite(value) or value < 0:
raise ProtocolError(ErrorCode.INVALID_ARGS,
'cutoff must be a finite, non-negative number of seconds')
return cls(cutoff=float(value))
CommandArgs = Union[HelloArgs, OnDemandStartArgs, OnDemandStopArgs, NoArgs,
BrightnessSetArgs, PluginReloadArgs, StateGetArgs, StateSubscribeArgs,
ErrorsClearArgs]
BrightnessSetArgs, PluginReloadArgs, StateGetArgs, StateSubscribeArgs]
#: The arguments of a command that goes on the render thread's queue.
QueuedArgs = Union[OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs]
@@ -613,7 +581,6 @@ _ARG_TYPES: Dict[str, Any] = {
Command.PLUGIN_RELOAD: PluginReloadArgs,
Command.STATE_GET: StateGetArgs,
Command.STATE_SUBSCRIBE: StateSubscribeArgs,
Command.ERRORS_CLEAR: ErrorsClearArgs,
}
@@ -686,13 +653,6 @@ class PluginReloadResult(TypedDict):
modes: List[str]
class ErrorsClearResult(TypedDict):
"""``errors.clear``, once applied and the error snapshot republished."""
request_id: str
cutoff: float
cleared: int
class LoopState(TypedDict):
"""``loop``: is the render loop still going round?
+5 -42
View File
@@ -6,9 +6,7 @@ rendering: a command that changes the panel is validated, put on a bounded
queue and acknowledged, and the render thread drains that queue at the point
where it reads the file mailbox (``DisplayController._poll_on_demand_requests``),
handing each command to the same code. Queries (``on_demand.status``) are
answered from a snapshot callable the display provides, and the few commands
that touch nothing the render thread owns (``errors.clear``) by a handler the
display registers, on the connection thread.
answered from a snapshot callable the display provides.
The queue also wakes the render thread: :meth:`ControlServer.wait_for_command`
is what it waits on in place of a sleep, so a command lands within a frame on
@@ -64,7 +62,6 @@ from src.ipc.contract import (
AWAITED_COMMANDS,
COMMANDS,
DEFAULT_SOCKET_DIR,
DIRECT_COMMANDS,
DEFAULT_SOCKET_PATH,
MAX_MESSAGE_BYTES,
MAX_SUBSCRIBERS,
@@ -110,8 +107,7 @@ MAX_CLIENTS = 8
#: Commands waiting for the render thread. It drains them at least every
#: 0.25 s, so a full queue means the render thread is stuck, and the client
#: is told ``busy`` instead of piling up work. The mailbox would not be read
#: either, so the web interface reports the failure rather than fall back.
#: is told ``busy`` (and falls back to the mailbox) instead of piling up work.
QUEUE_SIZE = 16
#: Timeout for one recv()/send() on a connection.
@@ -556,12 +552,6 @@ def server_socket_path(environ: Optional[Mapping[str, str]] = None) -> Optional[
StatusProvider = Callable[[], Dict[str, Any]]
#: A handler for one of DIRECT_COMMANDS, ``(request_id, args) -> result``. It
#: runs on the connection thread, so it must not touch what the render thread
#: owns. It may raise ProtocolError to answer with that error's code; any
#: other exception is answered ``internal``.
DirectHandler = Callable[[str, Any], Mapping[str, Any]]
class ControlServer:
"""Serves the control socket on background threads.
@@ -579,12 +569,9 @@ class ControlServer:
await_seconds: Optional[Mapping[str, float]] = None,
state_hub: Optional[StateHub] = None,
max_subscribers: int = MAX_SUBSCRIBERS,
keepalive: float = SUBSCRIBE_KEEPALIVE_SECONDS,
handlers: Optional[Mapping[str, DirectHandler]] = None):
keepalive: float = SUBSCRIBE_KEEPALIVE_SECONDS):
self.path = path
self.state_hub = state_hub
self._handlers: Dict[str, DirectHandler] = {
cmd: fn for cmd, fn in (handlers or {}).items() if cmd in DIRECT_COMMANDS}
self._subscriber_slots = threading.BoundedSemaphore(max_subscribers)
self._keepalive = keepalive
self._await_seconds: Dict[str, float] = dict(AWAIT_SECONDS)
@@ -1041,9 +1028,6 @@ class ControlServer:
snap = hub.snapshot()
return Response.success(request.id, fit_snapshot(snap), v=request.v)
if request.cmd in DIRECT_COMMANDS:
return self._direct(request, args)
if request.cmd in QUEUED_COMMANDS and isinstance(args, (
OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs)):
awaited = request.cmd in AWAITED_COMMANDS
@@ -1071,25 +1055,6 @@ class ControlServer:
return Response.failure(request.id, ErrorCode.INTERNAL,
f'{request.cmd} is not implemented', v=request.v)
def _direct(self, request: Request, args: Any) -> Response:
"""A command the display answers on this thread (DIRECT_COMMANDS)."""
handler = self._handlers.get(request.cmd)
if handler is None:
# Answered as an older display would, so the client falls back.
return Response.failure(request.id, ErrorCode.UNKNOWN_COMMAND,
f'{request.cmd} is not served by this display',
v=request.v)
try:
result = handler(request.id, args)
except ProtocolError as e:
return Response.failure(request.id, e.code, e.message, v=request.v)
except Exception: # pylint: disable=broad-except
logger.exception("Control socket: %s %s failed", request.cmd, request.id)
return Response.failure(request.id, ErrorCode.INTERNAL,
'the display failed to apply it', v=request.v)
logger.info("Control socket applied %s %s", request.cmd, request.id)
return Response.success(request.id, dict(result), v=request.v)
def _await_outcome(self, request: Request, outcome: CommandOutcome) -> Response:
"""Answer an awaited command once the render thread has applied it.
@@ -1114,9 +1079,7 @@ class ControlServer:
def start_control_server(status_provider: Optional[StatusProvider] = None,
cache_dir: Optional[str] = None,
environ: Optional[Mapping[str, str]] = None,
state_hub: Optional[StateHub] = None,
handlers: Optional[Mapping[str, DirectHandler]] = None,
) -> Optional[ControlServer]:
state_hub: Optional[StateHub] = None) -> Optional[ControlServer]:
"""Start the display's control socket, or return None when it can't run.
None covers Windows, ``LEDMATRIX_CONTROL_SOCKET=off`` and any failure to
@@ -1128,7 +1091,7 @@ def start_control_server(status_provider: Optional[StatusProvider] = None,
logger.debug("Control socket disabled or unsupported here; using the file mailbox only")
return None
server = ControlServer(path, status_provider, resolve_socket_group(cache_dir),
state_hub=state_hub, handlers=handlers)
state_hub=state_hub)
return server if server.start() else None
-128
View File
@@ -1,128 +0,0 @@
"""Keep glibc's malloc from holding on to memory the display has freed.
The display process allocates and frees PIL images and numpy buffers all day
from a dozen threads. glibc gives each allocating thread its own malloc arena
(up to 8 x CPU count) and returns little of what is freed inside them to the
OS, so resident memory climbs for hours while the live data stays flat. Two
in-process remedies, both standard library only (ctypes) and both no-ops off
Linux/glibc:
* :func:`cap_arenas` -- ``mallopt(M_ARENA_MAX, 2)``, the in-process twin of the
unit's ``Environment=MALLOC_ARENA_MAX=2``. Units installed before that line
existed never got it (systemd runs the copy in /etc/systemd/system), so the
process applies it itself. Call it before any other thread starts: arenas
already created stay. A ``MALLOC_ARENA_MAX`` set in the environment wins.
* :class:`MallocTrimmer` -- ``malloc_trim(0)`` at most every few minutes,
called from the render loop between screens, where no frame is being drawn.
glibc 2.8+ releases free pages from the middle of every arena, not only the
top of the main heap.
Without glibc (macOS, Windows, musl, the dev server on any of them) nothing is
loaded and every call returns False.
"""
import ctypes
import logging
import os
import sys
import time
from typing import Any, Callable, Optional
logger = logging.getLogger(__name__)
#: glibc's mallopt() parameter number for the arena cap (malloc.h).
M_ARENA_MAX = -8
#: The arena cap applied when the environment does not set one; the same value
#: as the unit's ``MALLOC_ARENA_MAX``.
DEFAULT_ARENA_MAX = 2
#: Seconds between malloc_trim() calls. A trim takes about 1-20 ms on a Pi 4,
#: so this keeps it far from frame timing while still returning memory long
#: before it piles up.
TRIM_INTERVAL_SECONDS = 300.0
_UNLOADED = object()
_libc: Any = _UNLOADED
def _load_libc() -> Optional[Any]:
"""The process's C library if it is glibc with malloc_trim, else None."""
global _libc
if _libc is _UNLOADED:
_libc = None
if sys.platform.startswith('linux'):
try:
libc = ctypes.CDLL(None)
# gnu_get_libc_version is glibc-only: musl also lacks
# malloc_trim, but this says why without guessing.
libc.gnu_get_libc_version
libc.malloc_trim.argtypes = [ctypes.c_size_t]
libc.malloc_trim.restype = ctypes.c_int
libc.mallopt.argtypes = [ctypes.c_int, ctypes.c_int]
libc.mallopt.restype = ctypes.c_int
_libc = libc
except (OSError, AttributeError, TypeError):
logger.debug("glibc malloc controls unavailable", exc_info=True)
return _libc
def cap_arenas(max_arenas: int = DEFAULT_ARENA_MAX) -> bool:
"""Cap glibc's malloc arenas at ``max_arenas``. True when the cap was set.
Skipped when ``MALLOC_ARENA_MAX`` is in the environment: glibc has read it
already, and an operator who set it chose that value.
"""
if os.environ.get('MALLOC_ARENA_MAX'):
return False
libc = _load_libc()
if libc is None:
return False
try:
return bool(libc.mallopt(M_ARENA_MAX, int(max_arenas)))
except Exception: # pylint: disable=broad-except
logger.debug("mallopt(M_ARENA_MAX) failed", exc_info=True)
return False
class MallocTrimmer:
"""Calls ``malloc_trim(0)`` at most once per ``interval`` seconds.
:meth:`maybe_trim` is meant for an idle point of the render loop; it costs
one clock read when no trim is due. The first trim comes one interval
after construction, so start-up's allocations have settled.
"""
def __init__(self, interval: float = TRIM_INTERVAL_SECONDS,
clock: Callable[[], float] = time.monotonic) -> None:
self._interval = interval
self._clock = clock
self._libc = _load_libc()
self._next = clock() + interval
@property
def available(self) -> bool:
return self._libc is not None
def maybe_trim(self) -> bool:
"""Trim if one is due. True when malloc_trim ran and released memory."""
if self._libc is None:
return False
now = self._clock()
if now < self._next:
return False
self._next = now + self._interval
try:
released = bool(self._libc.malloc_trim(0))
except Exception: # pylint: disable=broad-except
logger.debug("malloc_trim failed; not trying again", exc_info=True)
self._libc = None
return False
def _rss():
try:
with open('/proc/self/statm') as f:
return int(f.read().split()[1]) * 4096 // 1024
except Exception:
return -1
logger.info("EXP malloc_trim(0) took %.1f ms, released=%s rss_kb_after=%d",
(self._clock() - now) * 1000.0, released, _rss())
return released
-5
View File
@@ -138,11 +138,6 @@ class PluginStoreManager(_RegistryMixin, _InstallMixin, _UpdateMixin):
# the registry cache expires. Only one thread fetches; others wait and
# then get the result from the warm cache (double-checked locking).
self._registry_fetch_lock = threading.Lock()
# refresh_registry_in_background: the one refresh thread, and when
# an offline one may be retried (see that method).
self._registry_refresh_lock = threading.Lock()
self._registry_refresh_thread: Optional[threading.Thread] = None
self._registry_refresh_retry_after = 0.0
# Per-plugin locks for _reinstall_with_rollback: the web UI runs
# Flask with threaded=True, so two overlapping requests for the
+3 -57
View File
@@ -7,7 +7,6 @@ methods reach shared state and helpers through ``self``.
import json
import requests
import threading
import time
from concurrent.futures import ThreadPoolExecutor
from datetime import datetime
@@ -986,14 +985,10 @@ class _RegistryMixin:
def get_registry_info(self, plugin_id: str) -> Optional[Dict]:
"""
Get plugin information from the registry (plugins.json).
Get plugin information from the registry cache only (no GitHub API calls).
Makes no GitHub API calls, but it does go through `fetch_registry`:
when the in-memory copy is missing or older than
``registry_cache_timeout`` it downloads plugins.json, and with no
network that waits out the timeout and retries. A caller that must
not block on the network (the installed-plugins list) uses
`get_cached_registry_info` instead.
Use this for lightweight lookups where only registry fields are needed
(e.g., verified status, latest_version).
Args:
plugin_id: Plugin identifier
@@ -1004,52 +999,3 @@ class _RegistryMixin:
registry = self.fetch_registry()
plugins = registry.get('plugins', []) or []
return self._match_registry_entry(plugins, plugin_id)
def get_cached_registry_info(self, plugin_id: str) -> Optional[Dict]:
"""The registry entry for ``plugin_id`` from the copy already in
memory, however old; never touches the network.
None when no registry has been loaded yet, or the plugin isn't in it.
When the copy is missing or past ``registry_cache_timeout`` this
starts `refresh_registry_in_background`, so a later call has it.
"""
cache = getattr(self, 'registry_cache', None)
cache_time = getattr(self, 'registry_cache_time', None)
if (not cache or not cache_time
or (time.time() - cache_time) >= self.registry_cache_timeout):
self.refresh_registry_in_background()
plugins = cache.get('plugins') if isinstance(cache, dict) else None
if not isinstance(plugins, list):
return None
return self._match_registry_entry(
[p for p in plugins if isinstance(p, dict)], plugin_id)
def refresh_registry_in_background(self) -> bool:
"""Fetch the registry on a daemon thread; True when one was started.
At most one runs at a time. After a fetch that left no registry in
memory (offline), no new one starts for ``_failure_backoff_seconds``,
so an offline Pi doesn't retry on every page load.
"""
with self._registry_refresh_lock:
running = self._registry_refresh_thread
if running is not None and running.is_alive():
return False
if time.time() < self._registry_refresh_retry_after:
return False
thread = threading.Thread(
target=self._background_registry_refresh,
name='registry-refresh', daemon=True)
self._registry_refresh_thread = thread
thread.start()
return True
def _background_registry_refresh(self) -> None:
try:
self.fetch_registry()
except Exception as e: # noqa: BLE001 - a background warm-up must not crash
self.logger.warning("Background registry refresh failed: %s", e)
if not getattr(self, 'registry_cache', None):
with self._registry_refresh_lock:
self._registry_refresh_retry_after = (
time.time() + self._failure_backoff_seconds)
-484
View File
@@ -1,484 +0,0 @@
"""Runs one screen: the ScreenRunner of docs/RUN_LOOP_REDESIGN.md.
``ScreenRunner.run(plan, plugin)`` draws a screen's first frame, runs the
frame loop its plan's ``frame_policy`` picks (125 Hz or 1 Hz), makes up the
minimum duration when the loop ended early, and returns one
:class:`Outcome` saying why the screen ended. Everything that touches the
plugin, the panel or the controller's state goes through a
:class:`ScreenHost` (the DisplayController); everything that reads or waits
on the clock goes through an injected :class:`FrameClock`. The runner itself
holds no state between screens.
What can end a screen early is decided at the runner's service points: after
each frame, after the frame loop, and after the make-up dwell. At each one
the host gathers a snapshot and asks the Arbiter, once, whether a Source in
``plan.preemptible_by`` now wants the panel (:meth:`ScreenHost.check`). A yes
is ``ExitReason.PREEMPTED``: the next pass of the loop decides what shows,
and the rotation does not advance past the screen that was cut short.
The frame pacing is the loop that used to be inline in
``DisplayController.run()``, unchanged: the 125 Hz loop paces to an 8 ms
deadline from the start of each frame (sleeping at least 1 ms, so a frame
that overran still yields the GIL), and the 1 Hz loop sleeps a flat second
between frames, woken early by a control socket command.
"""
import logging
from dataclasses import dataclass
from enum import Enum
from typing import Any, NamedTuple, Optional, Protocol, Tuple
from src.display_arbiter import FramePolicy, ScreenPlan, Source
__all__ = [
"AFTER_COMPLETED_LOOP",
"AFTER_LOOP",
"Checkpoint",
"DYNAMIC_GRACE",
"ExitReason",
"FINAL",
"FRAME",
"FirstFrame",
"FrameClock",
"HIGH_FPS_INTERVAL",
"NoticeRead",
"Outcome",
"STATIC_INTERVAL",
"Screen",
"ScreenHost",
"ScreenRunner",
"after_dwell",
]
#: Seconds between frames in the high-FPS loop (125 Hz), for scrolling plugins.
HIGH_FPS_INTERVAL = 0.008
#: Seconds between frames in the static loop (1 Hz).
STATIC_INTERVAL = 1.0
#: A dynamic-duration screen ends on cycle completion only this long after its
#: minimum, so timing jitter around the minimum can't end it early.
DYNAMIC_GRACE = 0.5
class ExitReason(Enum):
"""Why a screen ended. The value is the golden traces' exit column where
one exists (test/test_run_loop_golden.py)."""
#: The screen ran its target duration.
DURATION = "duration"
#: A dynamic-duration plugin finished its cycle after its minimum.
CYCLE_COMPLETE = "cycle-complete"
#: The first frame had nothing to show (display() returned False or
#: raised inside the executor), or no plugin draws the mode.
EMPTY = "empty"
#: The first frame's dispatch itself raised.
ERROR = "error"
#: A later frame returned False (a dynamic-duration screen on the 1 Hz
#: loop keeps going instead).
DISPLAY_FALSE = "display-false"
#: Another Source took the panel, or an on-demand session ran out before
#: the screen began. The rotation does not advance.
PREEMPTED = "preempted"
#: A plugin reload is waiting for the top of the loop. The screen is cut
#: short but counts as shown: the rotation advances, and the next pass
#: reloads before it draws.
RELOAD = "reload"
@dataclass(frozen=True)
class Outcome:
"""How a screen ended.
Attributes:
exit_reason: Why it ended.
elapsed: Seconds from the end of the first frame to the end.
preempted_by: For PREEMPTED, the plan that took the panel when the
Arbiter named one (None when the session simply ran out).
on_demand_active: Filled in by the controller when the screen is
over: an on-demand session was running at that moment.
still_live: Filled in by the controller: the mode's plugin still had
live content at that moment, which holds the rotation on it.
"""
exit_reason: ExitReason
elapsed: float = 0.0
preempted_by: Optional[ScreenPlan] = None
on_demand_active: bool = False
still_live: bool = False
class FrameClock(Protocol):
"""The clocks the runner reads and the sleep it paces with.
The shape of the ``time`` module, so production passes it (through an
indirection that lets tests patch the module) and the golden traces pass
their fake clock.
"""
def time(self) -> float:
"""Wall-clock seconds: what screen durations are measured in."""
def perf_counter(self) -> float:
"""A monotonic high-resolution clock: what the 8 ms pacing reads."""
def sleep(self, seconds: float) -> None:
"""Block for ``seconds``."""
class NoticeRead(Enum):
"""When a service point reads the WiFi notice file.
The read is throttled to once a second and deletes an expired file, so
*when* it happens is behaviour: each service point reads it exactly
when the loop always did.
"""
#: Not at all.
NEVER = "never"
#: Only if nothing cheaper has already ended the screen: no on-demand
#: session, the panel on, the mode unchanged and no live takeover.
IF_UNDECIDED = "if-undecided"
#: Whenever no on-demand session is running.
ALWAYS = "always"
@dataclass(frozen=True)
class Checkpoint:
"""What one kind of service point considers.
Attributes:
name: For logs and tests.
notice: When the WiFi notice is read (see NoticeRead).
notice_counts: Whether a pending notice ends the screen here. After
the make-up dwell it does only while the screen had time left.
reload: Whether a pending plugin reload ends the screen here. Only
between frames: once the frame loop is over the screen is too.
"""
name: str
notice: NoticeRead
reload: bool
notice_counts: bool = True
#: Between frames: the frame loops' check, and the socket wake in the 1 Hz wait.
FRAME = Checkpoint("frame", NoticeRead.IF_UNDECIDED, reload=True)
#: After a frame loop that ended early (display() returned False, a reload).
AFTER_LOOP = Checkpoint("after-loop", NoticeRead.IF_UNDECIDED, reload=False)
#: After a frame loop that ran its course: only a mode change or the schedule.
AFTER_COMPLETED_LOOP = Checkpoint("after-completed-loop", NoticeRead.NEVER, reload=False)
#: The last look before the rotation advances.
FINAL = Checkpoint("final", NoticeRead.NEVER, reload=False)
def after_dwell(time_left: bool) -> Checkpoint:
"""After the make-up dwell: a notice that cut it short ends the screen,
so the mode resumes after the notice instead of rotating past it."""
return Checkpoint("after-dwell", NoticeRead.ALWAYS, reload=False,
notice_counts=time_left)
class FirstFrame(NamedTuple):
"""What the first frame's dispatch returned (see _dispatch_first_frame)."""
shown: bool
raised: bool
accepts_display_mode: bool
@dataclass
class Screen:
"""One running screen: its completed plan, the plugin drawing it and
when it started. Mutable only in that the runner owns it."""
plan: ScreenPlan
plugin: Any
accepts_display_mode: bool
start: float
@property
def mode(self) -> Optional[str]:
return self.plan.mode
class ScreenHost(Protocol):
"""The controller's side of a screen. See DisplayController."""
def first_frame(self, plan: ScreenPlan, plugin: Any) -> FirstFrame:
"""Draw the first frame through the plugin executor."""
def complete_plan(self, plan: ScreenPlan, plugin: Any) -> Optional[ScreenPlan]:
"""The plan with the plugin's durations, dynamic flag and frame
policy, read after the first frame. None when an on-demand session
has no time left for it."""
def draw(self, screen: Screen) -> Any:
"""One later frame: what display() returned."""
def after_frame(self, screen: Screen) -> None:
"""After a frame that did not end the screen (the follower frame)."""
def tick(self) -> None:
"""Plugin updates that have come due (throttled)."""
def service(self, screen: Screen) -> Optional[Tuple[str, ...]]:
"""Apply pending changes (on-demand requests, schedule, brightness,
finished reloads). Returns the live modes when a live-priority scan
was due, else None."""
def wait_frame(self, interval: float, screen: Screen) -> Optional[ScreenPlan]:
"""The 1 Hz loop's sleep between frames. The plan that takes the
panel when a control socket command ended the screen, else None."""
def check(self, screen: Screen, checkpoint: Checkpoint,
live_scan: Optional[Tuple[str, ...]] = None) -> Optional[ScreenPlan]:
"""The service point: the plan that now takes the panel from this
screen, or None while it holds. One Arbiter.decide() call."""
def dwell(self, seconds: float) -> None:
"""Sleep up to ``seconds``, servicing changes; returns early on one."""
def cycle_complete(self, screen: Screen) -> bool:
"""The plugin's dynamic-duration cycle is complete."""
class ScreenRunner:
"""Runs one screen at a time for a ScreenHost. See the module docstring."""
def __init__(self, clock: FrameClock, host: ScreenHost,
log: Optional[logging.Logger] = None):
self.clock = clock
self.host = host
# The controller passes its own logger, so these lines keep the
# source they always had in the journal.
self.log = log or logging.getLogger(__name__)
# -- the screen ------------------------------------------------------
def run(self, plan: ScreenPlan, plugin: Any) -> Outcome:
"""Run ``plan``'s screen, drawn by ``plugin`` (None: nothing draws it)."""
if plugin is None:
return Outcome(ExitReason.EMPTY)
first = self.host.first_frame(plan, plugin)
if not first.shown:
return Outcome(ExitReason.ERROR if first.raised else ExitReason.EMPTY)
completed = self.host.complete_plan(plan, plugin)
if completed is None:
return Outcome(ExitReason.PREEMPTED)
screen = Screen(completed, plugin, first.accepts_display_mode,
start=self.clock.time())
if completed.frame_policy is FramePolicy.HIGH_FPS:
reason, by = self._high_fps_loop(screen)
else:
reason, by = self._static_loop(screen)
if reason is ExitReason.PREEMPTED:
# The service point that ended the loop has decided; looking
# again now, at the same instant, gives the same answer.
return self._outcome(screen, reason, by)
loop_completed = reason in (ExitReason.DURATION, ExitReason.CYCLE_COMPLETE)
# LOAD-BEARING: a change the frame loop did not end on (a dwell
# inside it, a later frame returning False) must not fall into the
# make-up dwell below. It can run for the rest of the screen's
# duration, and a freshly requested on-demand mode would sit
# invisible for that long -- or be clobbered by a queued stop.
by = self.host.check(screen, AFTER_COMPLETED_LOOP if loop_completed else AFTER_LOOP)
if by is not None:
return self._outcome(screen, ExitReason.PREEMPTED, by)
# Honour the minimum duration when a static, non-dynamic screen's
# loop ended early. A screen cut short for a plugin reload is over:
# the dwell returns at once and the rotation advances.
if (not completed.dynamic and not loop_completed
and completed.frame_policy is not FramePolicy.HIGH_FPS):
elapsed = self.clock.time() - screen.start
remaining = max(0.0, self._max(screen) - elapsed)
if remaining > 0:
self.host.dwell(remaining)
time_left = self.clock.time() - screen.start < self._max(screen)
by = self.host.check(screen, after_dwell(time_left))
if by is not None:
return self._outcome(screen, ExitReason.PREEMPTED, by)
if completed.dynamic:
self._log_dynamic_end(screen)
# The dwells above return early when a pending change (on-demand
# started or stopped, the panel scheduled off) has already decided
# what comes next; rotating now would skip it.
by = self.host.check(screen, FINAL)
if by is not None:
return self._outcome(screen, ExitReason.PREEMPTED, by)
return self._outcome(screen, reason, None)
def _outcome(self, screen: Screen, reason: ExitReason,
by: Optional[ScreenPlan]) -> Outcome:
return Outcome(reason, self.clock.time() - screen.start, preempted_by=by)
@staticmethod
def _max(screen: Screen) -> float:
return float(screen.plan.max_duration or 0.0)
@staticmethod
def _min(screen: Screen) -> float:
return float(screen.plan.min_duration or 0.0)
@staticmethod
def _ended_by(by: ScreenPlan) -> ExitReason:
return ExitReason.RELOAD if by.source is Source.RELOAD else ExitReason.PREEMPTED
# -- the frame loops ---------------------------------------------------
def _high_fps_loop(self, screen: Screen) -> Tuple[ExitReason, Optional[ScreenPlan]]:
"""Ultra-smooth frames for scrolling plugins (8 ms = 125 FPS)."""
clock, host, log = self.clock, self.host, self.log
interval = HIGH_FPS_INTERVAL
log.debug("Entering high-FPS loop for %s with display_interval=%.3fs (%.1f FPS)",
screen.mode, interval, 1.0 / interval)
target = self._max(screen)
while True:
frame_start = clock.perf_counter()
try:
result = host.draw(screen)
if isinstance(result, bool) and not result:
log.debug("Display returned False, breaking early")
return ExitReason.DISPLAY_FALSE, None
except Exception: # pylint: disable=broad-except
log.exception("Error during display update")
# Multi-display sync: send follower frame after each render
host.after_frame(screen)
host.tick()
# Throttled: one clock compare between passes. A live-priority
# scan, when one is due, happens here, before the sleep, as it
# always has; the Arbiter weighs it after the sleep.
live_scan = host.service(screen)
# Pace to the frame deadline rather than sleeping a flat
# interval on top of the work. display() has already blocked on
# the panel's vsync by this point, so an unconditional sleep is
# added to a wait that already happened. Measured on a 2x128x64
# chain at limit_refresh_rate_hz=100: ~4ms of render plus a flat
# 8ms put each iteration at ~12ms against a 10ms refresh grid, so
# every swap missed a refresh and the loop settled at 50fps where
# display_interval asks for 125 -- and with zero headroom, ~14% of
# frames slipped a further refresh, which is what reads as scroll
# stutter.
remaining = interval - (clock.perf_counter() - frame_start)
# Yield even when the frame overran its budget, so plugin update
# threads and the web UI are not starved of the GIL.
clock.sleep(remaining if remaining > 0 else 0.001)
by = host.check(screen, FRAME, live_scan)
if by is not None:
log.debug("Mode changed during high-FPS loop, breaking early")
return self._ended_by(by), by
elapsed = clock.time() - screen.start
if elapsed >= target:
log.debug("Reached high-FPS target duration %.2fs for mode %s",
target, screen.mode)
return ExitReason.DURATION, None
if self._should_exit_dynamic(screen, elapsed):
log.debug("Dynamic duration cycle complete for %s after %.2fs",
screen.mode, elapsed)
return ExitReason.CYCLE_COMPLETE, None
def _static_loop(self, screen: Screen) -> Tuple[ExitReason, Optional[ScreenPlan]]:
"""One frame a second for everything else."""
clock, host, log = self.clock, self.host, self.log
interval = STATIC_INTERVAL
log.debug("Entering normal FPS loop for %s with display_interval=%.3fs",
screen.mode, interval)
target = self._max(screen)
dynamic = screen.plan.dynamic
while True:
# Wakes for a control socket command and applies it at once,
# instead of up to a second later.
by = host.wait_frame(interval, screen)
if by is not None:
log.info("Mode changed during display loop from %s to %s (%s), "
"breaking early", screen.mode, by.mode, by.source.value)
return self._ended_by(by), by
host.tick()
elapsed = clock.time() - screen.start
if elapsed >= target:
log.debug("Reached standard target duration %.2fs for mode %s",
target, screen.mode)
return ExitReason.DURATION, None
try:
result = host.draw(screen)
if isinstance(result, bool) and not result:
# A dynamic-duration screen doesn't end on False: it
# keeps looping until its cycle completes or its maximum.
if not dynamic:
log.info("Display returned False for %s (no dynamic duration), "
"breaking early", screen.mode)
return ExitReason.DISPLAY_FALSE, None
log.debug("Display returned False for %s (dynamic duration enabled), "
"continuing loop", screen.mode)
except Exception: # pylint: disable=broad-except
log.exception("Error during display update")
# Multi-display sync: send follower frame after each render
host.after_frame(screen)
live_scan = host.service(screen)
by = host.check(screen, FRAME, live_scan)
if by is not None:
log.info("Mode changed during display loop from %s to %s (%s), "
"breaking early", screen.mode, by.mode, by.source.value)
return self._ended_by(by), by
if self._should_exit_dynamic(screen, elapsed):
log.info("Dynamic duration cycle complete for %s after %.2fs",
screen.mode, elapsed)
return ExitReason.CYCLE_COMPLETE, None
# -- dynamic duration --------------------------------------------------
def _should_exit_dynamic(self, screen: Screen, elapsed: float) -> bool:
if not screen.plan.dynamic:
return False
minimum = self._min(screen)
# A small grace period after min_duration prevents premature exits
# due to timing issues.
if elapsed < minimum + DYNAMIC_GRACE:
self.log.debug(
"_should_exit_dynamic: elapsed %.2fs < min_duration %.2fs + grace %.2fs, "
"returning False", elapsed, minimum, DYNAMIC_GRACE)
return False
cycle_complete = self.host.cycle_complete(screen)
self.log.debug(
"_should_exit_dynamic: elapsed %.2fs >= min %.2fs, cycle_complete=%s, returning %s",
elapsed, minimum + DYNAMIC_GRACE, cycle_complete, cycle_complete)
if cycle_complete:
self.log.debug("Cycle complete detected for %s after %.2fs (min: %.2fs, grace: %.2fs)",
screen.mode, elapsed, minimum, DYNAMIC_GRACE)
return cycle_complete
def _log_dynamic_end(self, screen: Screen) -> None:
"""How a dynamic-duration screen ended, for the log. Asks the plugin
once more whether its cycle is complete, as the loop always did."""
elapsed_total = self.clock.time() - screen.start
cycle_done = self.host.cycle_complete(screen)
minimum, maximum = self._min(screen), self._max(screen)
if cycle_done:
self.log.info(
"Dynamic duration cycle completed for %s after %.2fs "
"(target: %.2fs, min: %.2fs, max: %.2fs)",
screen.mode, elapsed_total, maximum, minimum, maximum)
elif elapsed_total >= maximum:
self.log.info(
"Dynamic duration cap reached before cycle completion for %s "
"(%.2fs/%ds, min: %.2fs)",
screen.mode, elapsed_total, int(maximum), minimum)
else:
self.log.debug(
"Dynamic duration cycle in progress for %s: %.2fs elapsed "
"(target: %.2fs, min: %.2fs, max: %.2fs)",
screen.mode, elapsed_total, maximum, minimum, maximum)
-12
View File
@@ -143,24 +143,12 @@ class FakeCache:
def __init__(self):
self.data: Dict[str, Any] = {}
self.cache_dir = "/nonexistent/run-loop-harness"
self._writes = 0
self._written: Dict[str, int] = {}
def get(self, key, max_age=None, memory_ttl=None):
return self.data.get(key)
def set(self, key, data, ttl=None):
self.data[key] = data
# Every write is a new file, as DiskCache's rename makes it.
self._writes += 1
self._written[key] = self._writes
def file_signature(self, key):
"""CacheManager.file_signature: None without a file, else a value
that changes with every write."""
if key not in self.data:
return None
return (self._written.get(key, 0), 0, 0)
def delete(self, key):
self.data.pop(key, None)
+9
View File
@@ -175,6 +175,15 @@
"POST"
]
],
[
"/api/v3/config/refresh-rate",
"api_v3.get_refresh_rate",
[
"GET",
"HEAD",
"OPTIONS"
]
],
[
"/api/v3/config/schedule",
"api_v3.get_schedule_config",
+1 -1
View File
@@ -29,7 +29,7 @@ def installed(api_v3_module, api_v3_client, tmp_path):
api.plugin_catalog.plugins_dir = str(tmp_path) # no manifest on disk
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[info])
api.plugin_catalog.get_plugin_display_modes = MagicMock(return_value=declared_modes)
api.plugin_store_manager.get_cached_registry_info = MagicMock(return_value=None)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={})
response = api_v3_client.get('/api/v3/plugins/installed')
assert response.status_code == 200
+1 -1
View File
@@ -22,7 +22,7 @@ def installed(api_v3_module, api_v3_client, tmp_path):
info.update(manifest_extra)
api.plugin_catalog.plugins_dir = str(tmp_path) # no manifest on disk
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[info])
api.plugin_store_manager.get_cached_registry_info = MagicMock(return_value=None)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={})
response = api_v3_client.get('/api/v3/plugins/installed')
assert response.status_code == 200
+7 -87
View File
@@ -1,13 +1,10 @@
"""POST /display/on-demand/start and /stop: control socket first, mailbox fallback.
The routes hand the request to the display over the control socket
(src/ipc) and get an acknowledgement. When the socket could not carry the
request -- no socket (a stopped display, or one older than the socket), a
refused or timed-out connect, a display too old to know the command, a bug
in the client -- they write the file mailbox exactly as they did before the
socket existed. When the display had the request and failed it (a full
queue, bad arguments, no answer in time) the route says so and writes
nothing. These tests pin each path, that at most one of them is used, that
(src/ipc) and get an acknowledgement. On any failure -- no socket (a stopped
display, or one older than the socket), a timeout, a refusal, a bug in the
client -- they write the file mailbox exactly as they did before the socket
existed. These tests pin both paths, that exactly one of them is used, that
the response says which, and that the request id is the same either way (the
display deduplicates on it).
@@ -104,15 +101,12 @@ class TestSocketPath:
class TestMailboxFallback:
@pytest.mark.parametrize("reason", [
"no_socket", "refused", "timeout", "closed", "bad_response", "invalid_request",
"busy", "forbidden", "unknown_command", "unsupported_version", "disabled",
"unsupported",
"busy", "unknown_command", "unsupported_version", "disabled", "unsupported",
])
def test_a_request_the_socket_never_carried_writes_the_mailbox(
def test_any_socket_failure_writes_the_mailbox_as_before(
self, api_v3_client, service, reason):
# sent=False: the display never had it (no socket, a refused or
# timed-out connect, turned away at the door).
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError(reason, "x", sent=False)):
side_effect=control_client.ControlError(reason, "x")):
resp = api_v3_client.post(START_URL, json={
"plugin_id": "weather", "mode": "weather_current",
"duration": 60, "pinned": True})
@@ -174,60 +168,6 @@ class TestMailboxFallback:
assert data["socket_error"] in ("disabled", "unsupported") # Linux, Windows
class TestTheDisplayHadIt:
"""Once the display has the request, its answer stands: no mailbox copy.
A busy queue, a refusal or silence after the request was sent mean the
display may have applied it, or would refuse the mailbox copy too, so
the route reports the failure instead of posting it a second time.
"""
@pytest.mark.parametrize("reason,status", [
("busy", 503), ("internal", 503), ("timeout", 503), ("closed", 503),
("bad_response", 503), ("invalid_args", 400),
])
def test_start_is_answered_with_the_failure(self, api_v3_client, service, reason, status):
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError(reason, "x", sent=True)):
resp = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert resp.status_code == status
body = resp.get_json()
assert body["status"] == "error"
assert body["data"]["transport"] == "socket"
assert body["data"]["socket_error"] == reason
assert _mailbox_writes(service["cache"]) == []
assert not [call for call in service["calls"] if call[0] == "systemctl"]
def test_stop_is_answered_with_the_failure(self, api_v3_client, service):
with patch(f"{CLIENT}.on_demand_stop",
side_effect=control_client.ControlError("busy", "x", sent=True)):
resp = api_v3_client.post(STOP_URL, json={})
assert resp.status_code == 503
assert _mailbox_writes(service["cache"]) == []
def test_a_stop_with_stop_service_still_stops_the_service(self, api_v3_client, service):
with patch(f"{CLIENT}.on_demand_stop",
side_effect=control_client.ControlError("timeout", "x", sent=True)), \
patch("web_interface.blueprints.api_v3.display._stop_display_service",
return_value={"active": False}) as stop:
resp = api_v3_client.post(STOP_URL, json={"stop_service": True})
assert resp.status_code == 200
data = resp.get_json()["data"]
assert data["transport"] == "socket" and data["socket_error"] == "timeout"
stop.assert_called_once()
assert _mailbox_writes(service["cache"]) == []
@pytest.mark.parametrize("reason", ["unknown_command", "unsupported_version"])
def test_an_older_display_that_does_not_speak_it_gets_the_mailbox(
self, api_v3_client, service, reason):
# The upgrade case: new web interface, display still on an old build.
with patch(f"{CLIENT}.on_demand_start",
side_effect=control_client.ControlError(reason, "x", sent=True)):
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
assert data["transport"] == "mailbox" and data["socket_error"] == reason
assert len(_mailbox_writes(service["cache"])) == 1
@pytest.mark.skipif(not c.socket_supported(), reason="AF_UNIX sockets are Linux/macOS only")
class TestRealSocket:
@pytest.fixture
@@ -264,23 +204,3 @@ class TestRealSocket:
data = api_v3_client.post(START_URL, json={"plugin_id": "weather"}).get_json()["data"]
assert data["transport"] == "mailbox" and data["socket_error"] == "no_socket"
assert len(_mailbox_writes(service["cache"])) == 1
def test_a_full_queue_is_reported_not_mailed(self, api_v3_client, service, monkeypatch):
import shutil
import tempfile
from src.ipc.server import ControlServer
d = tempfile.mkdtemp(prefix="lmipc-")
path = os.path.join(d, "control.sock")
server = ControlServer(path, status_provider=dict, queue_size=1)
assert server.start()
monkeypatch.setenv(c.SOCKET_PATH_ENV, path)
try:
first = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert first.get_json()["data"]["transport"] == "socket"
second = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert second.status_code == 503
assert second.get_json()["data"]["socket_error"] == "busy"
assert _mailbox_writes(service["cache"]) == []
finally:
server.close()
shutil.rmtree(d, ignore_errors=True)
+12 -399
View File
@@ -9,7 +9,6 @@ that overrides the schedule and ends (#714).
"""
import itertools
from dataclasses import replace
import os
from unittest.mock import MagicMock, patch
@@ -20,7 +19,6 @@ os.environ.setdefault("EMULATOR", "true")
from src import display_arbiter # noqa: E402
from src.display_arbiter import ( # noqa: E402
SCHEDULED_OFF_DWELL,
SCREEN_PREEMPTERS,
WIFI_NOTICE_DWELL,
Arbiter,
ArbiterInputs,
@@ -28,9 +26,6 @@ from src.display_arbiter import ( # noqa: E402
ScreenPlan,
Source,
WifiNotice,
live_pick,
live_takeover,
on_demand_bound,
wifi_notice_preempts,
)
@@ -39,9 +34,7 @@ NOTICE = WifiNotice(message="Connected to HomeNet", expires_at=1_000.0)
OFF = Source.SCHEDULED_OFF
FOLLOW = Source.FOLLOWER
WIFI = Source.WIFI
ONDEM = Source.ON_DEMAND
LEGACY = Source.LEGACY
ROTATION = Source.ROTATION
# (schedule_on, on_demand_active, follower_active, notice) -> Source.
# Every combination of the stage-2 inputs: 2 x 2 x 2 x 2 = 16 rows.
@@ -53,17 +46,17 @@ DECIDE_TABLE = [
(False, False, True, None, OFF),
(False, False, True, NOTICE, OFF),
# Scheduled off, but on-demand overrides the gate.
(False, True, False, None, ONDEM), # on-demand overrides the gate
(False, True, False, NOTICE, ONDEM), # on-demand outranks WiFi
(False, True, False, None, LEGACY), # on-demand: run() decides
(False, True, False, NOTICE, LEGACY), # on-demand outranks WiFi
(False, True, True, None, FOLLOW), # follower outranks on-demand
(False, True, True, NOTICE, FOLLOW),
# Scheduled on.
(True, False, False, None, ROTATION), # live / Vegas / rotation
(True, False, False, None, LEGACY), # live / Vegas / rotation
(True, False, False, NOTICE, WIFI),
(True, False, True, None, FOLLOW),
(True, False, True, NOTICE, FOLLOW), # follower outranks WiFi
(True, True, False, None, ONDEM),
(True, True, False, NOTICE, ONDEM), # on-demand outranks WiFi
(True, True, False, None, LEGACY),
(True, True, False, NOTICE, LEGACY), # on-demand outranks WiFi
(True, True, True, None, FOLLOW),
(True, True, True, NOTICE, FOLLOW),
]
@@ -102,15 +95,8 @@ class TestDecide:
elif expected is WIFI:
assert plan == ScreenPlan(WIFI, max_duration=WIFI_NOTICE_DWELL,
notice=NOTICE)
elif expected is ONDEM:
# An empty state: a session with no modes, which the
# controller ends (see TestOnDemand for real sessions).
assert plan == ScreenPlan(ONDEM)
elif expected is ROTATION:
# No live scan and Vegas off: the rotation's (empty) mode.
assert plan == ScreenPlan(ROTATION, preemptible_by=SCREEN_PREEMPTERS)
else:
# A follower paces itself.
# A follower paces itself; LEGACY is run()'s existing code.
assert plan == ScreenPlan(expected)
def test_dwells_are_todays(self):
@@ -238,389 +224,16 @@ class TestControllerSnapshot:
assert (plan.source is OFF) is (not dc.is_display_active), time_str
return plan.source
assert step("22:59:30") is ROTATION
assert step("22:59:30") is LEGACY
assert step("23:00:00") is OFF # window ends
dc.on_demand_active = True
assert step("23:00:10") is ONDEM # on-demand overrides
assert step("23:00:10") is LEGACY # on-demand overrides
assert dc.on_demand_schedule_override is True
assert step("23:01:00") is ONDEM # next minute, still on
assert step("23:01:00") is LEGACY # next minute, still on
dc._reset_on_demand_fields() # session ends
assert step("23:01:20") is OFF # same minute: blanks
dc.on_demand_active = True
assert step("06:59:00") is ONDEM
assert step("07:00:00") is ONDEM # schedule back on mid-session
assert step("06:59:00") is LEGACY
assert step("07:00:00") is LEGACY # schedule back on mid-session
dc._reset_on_demand_fields()
assert step("07:00:30") is ROTATION
# -- OnDemand (stage 3) ---------------------------------------------------
ON = ArbiterInputs(schedule_on=True, on_demand_active=True, follower_active=False)
def _session(modes=("a", "b", "c"), index=0, expires_at=None, current=None):
return ArbiterState(current_mode=current, on_demand_modes=tuple(modes),
on_demand_index=index, on_demand_expires_at=expires_at)
class TestOnDemand:
# (modes, index, expires_at, now) -> (mode, max_duration)
TABLE = [
(("a", "b", "c"), 0, None, 100.0, "a", None), # untimed
(("a", "b", "c"), 2, None, 100.0, "c", None),
(("a", "b", "c"), 3, None, 100.0, "a", None), # past the end: 0
(("a", "b", "c"), 9, None, 100.0, "a", None),
(("a",), 0, 130.0, 100.0, "a", 30.0), # 30 s left
(("a",), 0, 130.0, 130.0, "a", 0.0), # none left
(("a",), 0, 130.0, 200.0, "a", 0.0), # never negative
]
@pytest.mark.parametrize("modes,index,expires_at,now,mode,max_duration", TABLE)
def test_current_mode_and_time_left(self, modes, index, expires_at, now, mode,
max_duration):
plan = Arbiter.decide(_session(modes, index, expires_at), ON, now)
assert plan.source is ONDEM
assert plan.mode == mode
assert plan.max_duration == max_duration
assert plan.deadline == expires_at
def test_no_modes_left_is_a_plan_with_no_mode(self):
plan = Arbiter.decide(_session(modes=()), ON, 0.0)
assert plan == ScreenPlan(ONDEM)
def test_preemptible_by_the_schedule_a_reload_and_its_own_changes(self):
plan = Arbiter.decide(_session(), ON, 0.0)
assert {Source.SCHEDULED_OFF, ONDEM, Source.RELOAD} <= plan.preemptible_by
assert Source.FOLLOWER not in plan.preemptible_by
@pytest.mark.parametrize("index,expected_index,expected_mode", [
(0, 1, "b"), (1, 2, "c"), (2, 0, "a")])
def test_next_on_demand_wraps(self, index, expected_index, expected_mode):
nxt = _session(index=index).next_on_demand()
assert (nxt.on_demand_index, nxt.current_mode) == (expected_index, expected_mode)
def test_showing_a_plan_resets_an_index_past_the_end(self):
state = _session(index=5, current="x")
plan = Arbiter.decide(state, ON, 0.0)
shown = state.showing(plan)
assert (shown.on_demand_index, shown.current_mode) == (0, "a")
assert state.on_demand_index == 5 # not mutated
# (min, max, deadline, now) -> bounds. The bound applied after the first frame.
BOUND_TABLE = [
(10.0, 20.0, None, 0.0, (10.0, 20.0)), # untimed: unchanged
(10.0, 20.0, 100.0, 50.0, (10.0, 20.0)), # plenty left
(10.0, 20.0, 100.0, 85.0, (10.0, 15.0)), # max cut to what is left
(10.0, 20.0, 100.0, 95.0, (5.0, 5.0)), # both cut
(10.0, 20.0, 100.0, 100.0, None), # nothing left
(10.0, 20.0, 100.0, 150.0, None),
]
@pytest.mark.parametrize("min_d,max_d,deadline,now,expected", BOUND_TABLE)
def test_on_demand_bound(min_d, max_d, deadline, now, expected):
assert on_demand_bound(min_d, max_d, deadline, now) == expected
# -- Live (stage 3) -------------------------------------------------------
LIVE = Source.LIVE
# (live_modes, current_mode, advance) -> pick. _check_live_priority's rule.
PICK_TABLE = [
((), "clock", True, None),
(None, "clock", True, None),
(("nfl",), "clock", True, "nfl"), # not on a live mode: first
(("nfl", "nhl"), "clock", True, "nfl"),
(("nfl", "nhl"), "clock", False, "nfl"),
(("nfl", "nhl"), "nfl", True, "nhl"), # round-robin
(("nfl", "nhl"), "nhl", True, "nfl"), # wraps
(("nfl", "nhl"), "nhl", False, "nhl"), # a peek stays put
(("nfl",), "nfl", True, "nfl"), # one game: itself
]
@pytest.mark.parametrize("live,current,advance,expected", PICK_TABLE)
def test_live_pick(live, current, advance, expected):
assert live_pick(live, current, advance) == expected
def _below(live=None, vegas=False, keeps=False, yielded=False, on_demand=False):
"""Inputs for the Sources below the notice (no notice, no follower)."""
return ArbiterInputs(schedule_on=True, on_demand_active=on_demand,
follower_active=False, live_modes=live, vegas_enabled=vegas,
vegas_live_in_ticker=keeps, vegas_yielded=yielded)
ROT = ("clock", "weather", "nfl_live", "nhl_live")
def _rot(current="clock", index=0, resume=None, unshown=False):
return ArbiterState(current_mode=current, rotation=ROT, rotation_index=index,
live_resume_index=resume, live_takeover_unshown=unshown)
class TestLive:
# (state, inputs) -> (source, mode, ends_live)
TABLE = [
# Nothing live, nothing to resume.
(_rot(), _below(live=()), "below", None, False),
# Not scanned (on-demand, or the ticker keeps live content).
(_rot(), _below(live=None), "below", None, False),
# A game is live: it takes the panel.
(_rot(), _below(live=("nfl_live",)), LIVE, "nfl_live", False),
# Two: round-robin from the one showing.
(_rot("nfl_live", 2), _below(live=("nfl_live", "nhl_live")), LIVE, "nhl_live", False),
# ... unless a mid-screen takeover chose it and it has not shown yet.
(_rot("nfl_live", 2, resume=0, unshown=True),
_below(live=("nfl_live", "nhl_live")), LIVE, "nfl_live", False),
# Vegas keeps live content in its ticker: Live has no say at all,
# not even the resume.
(_rot(), _below(live=("nfl_live",), vegas=True, keeps=True), "below", None, False),
(_rot("nfl_live", 2, resume=1),
_below(live=(), vegas=True, keeps=True), "below", None, False),
# Vegas that yields to live content: Live outranks it.
(_rot(), _below(live=("nfl_live",), vegas=True), LIVE, "nfl_live", False),
# The game ended: the interrupted rotation resumes.
(_rot("nfl_live", 2, resume=1), _below(live=()), "below", None, True),
(_rot("nfl_live", 2, resume=1), _below(live=(), vegas=True), "below", None, True),
# On-demand outranks Live.
(ArbiterState(on_demand_modes=("x",)), _below(live=("nfl_live",), on_demand=True),
ONDEM, "x", False),
]
@pytest.mark.parametrize("state,inputs,source,mode,ends_live", TABLE)
def test_decide(self, state, inputs, source, mode, ends_live):
plan = Arbiter.decide(state, inputs, 0.0)
if source == "below":
assert plan.source not in (LIVE, ONDEM, OFF, FOLLOW, WIFI)
else:
assert plan.source is source
assert plan.mode == mode
assert plan.ends_live is ends_live
def test_a_live_plan_has_no_durations_until_its_first_frame(self):
plan = Arbiter.decide(_rot(), _below(live=("nfl_live",)), 0.0)
assert (plan.min_duration, plan.max_duration, plan.frame_policy) == (None, None, None)
class TestLiveTransitions:
def test_claim_saves_where_the_rotation_was(self):
nxt = _rot("weather", 1).claim_live("nfl_live")
assert (nxt.current_mode, nxt.rotation_index, nxt.live_resume_index) == ("nfl_live", 2, 1)
def test_a_second_claim_keeps_the_first_resume_point(self):
nxt = _rot("nfl_live", 2, resume=1).claim_live("nhl_live")
assert (nxt.current_mode, nxt.rotation_index, nxt.live_resume_index) == ("nhl_live", 3, 1)
def test_claiming_the_mode_showing_changes_nothing(self):
state = _rot("nfl_live", 2, resume=1)
assert state.claim_live("nfl_live") is state
def test_a_live_mode_outside_the_rotation_keeps_the_index(self):
nxt = _rot("weather", 1).claim_live("mlb_live")
assert (nxt.current_mode, nxt.rotation_index, nxt.live_resume_index) == ("mlb_live", 1, 1)
def test_release_resumes_and_forgets(self):
nxt = _rot("nhl_live", 3, resume=1).release_live()
assert (nxt.current_mode, nxt.rotation_index, nxt.live_resume_index) == ("weather", 1, None)
def test_release_wraps_a_resume_point_past_a_shortened_rotation(self):
nxt = _rot("nhl_live", 3, resume=6).release_live()
assert (nxt.current_mode, nxt.rotation_index) == ("nfl_live", 2) # 6 % 4
def test_release_with_nothing_to_resume_changes_nothing(self):
state = _rot("clock", 0)
assert state.release_live() is state
empty = ArbiterState(current_mode="x", live_resume_index=2)
assert empty.release_live() is empty # no rotation to resume into
# -- Vegas and Rotation (stage 3) -----------------------------------------
class TestVegasAndRotation:
# (state, inputs) -> (source, mode, ends_live). LEGACY now means Vegas only.
TABLE = [
(_rot("weather", 1), _below(live=()), ROTATION, "weather", False),
(_rot("weather", 1), _below(live=None), ROTATION, "weather", False),
(_rot("weather", 1), _below(live=(), vegas=True), LEGACY, None, False),
(_rot("weather", 1), _below(live=None, vegas=True, keeps=True), LEGACY, None, False),
# The iteration yielded: the screen it fell through to.
(_rot("weather", 1), _below(live=(), vegas=True, yielded=True),
ROTATION, "weather", False),
(_rot("weather", 1), _below(live=("nfl_live",), vegas=True, yielded=True),
LIVE, "nfl_live", False),
# Live priority just ended: the rotation resumes where it was cut.
(_rot("nhl_live", 3, resume=1), _below(live=()), ROTATION, "weather", True),
# ... and Vegas carries the resume through to its own pass.
(_rot("nhl_live", 3, resume=1), _below(live=(), vegas=True), LEGACY, None, True),
# A rotation that something moved off its list carries on from there.
(ArbiterState(current_mode=None, rotation=ROT), _below(live=()), ROTATION, None, False),
]
@pytest.mark.parametrize("state,inputs,source,mode,ends_live", TABLE)
def test_decide(self, state, inputs, source, mode, ends_live):
plan = Arbiter.decide(state, inputs, 0.0)
assert (plan.source, plan.mode, plan.ends_live) == (source, mode, ends_live)
def test_a_rotation_plan_may_be_preempted_by_everything_a_screen_watches(self):
plan = Arbiter.decide(_rot(), _below(live=()), 0.0)
assert plan.preemptible_by == SCREEN_PREEMPTERS
assert {OFF, ONDEM, WIFI, LIVE, ROTATION, Source.RELOAD} == SCREEN_PREEMPTERS
class _End:
def __init__(self, on_demand_active=False, still_live=False):
self.on_demand_active = on_demand_active
self.still_live = still_live
class TestAfter:
"""ArbiterState.after: _advance_after_screen's step."""
def test_the_rotation_advances(self):
nxt = _rot("weather", 1).after(_End())
assert (nxt.current_mode, nxt.rotation_index) == ("nfl_live", 2)
def test_it_wraps(self):
nxt = _rot("nhl_live", 3).after(_End())
assert (nxt.current_mode, nxt.rotation_index) == ("clock", 0)
def test_a_live_mode_still_live_holds(self):
state = _rot("nfl_live", 2)
assert state.after(_End(still_live=True)) is state
def test_an_on_demand_session_moves_to_its_next_mode(self):
state = replace(_session(index=1, current="b"), rotation=ROT, rotation_index=1)
nxt = state.after(_End(on_demand_active=True))
assert (nxt.current_mode, nxt.on_demand_index, nxt.rotation_index) == ("c", 2, 1)
def test_a_session_with_no_modes_is_left_to_the_controller(self):
state = ArbiterState(current_mode="x", rotation=ROT)
assert state.after(_End(on_demand_active=True)) is state
def test_no_rotation_no_step(self):
state = ArbiterState(current_mode="x")
assert state.after(_End()) is state
# -- Mid-screen: decide(..., running=plan) (stage 3) ----------------------
#
# What the ScreenRunner asks at each service point. These rows are what
# _check_live_takeover, _screen_preempted and _wifi_notice_pending answered
# between frames before stage 3, written out.
CLOCK = Arbiter.decide(ArbiterState(current_mode="clock"), _below(live=()), 0.0)
NFL = Arbiter.decide(ArbiterState(current_mode="clock"), _below(live=("nfl_live",)), 0.0)
OD = Arbiter.decide(_session(modes=("x", "y")), ON, 0.0)
FRESH = WifiNotice(message="AP mode", expires_at=1_000.0)
HELD = "held"
def _mid(on_demand=False, schedule_on=True, live=None, notice=None, reload=False):
return ArbiterInputs(schedule_on=schedule_on, on_demand_active=on_demand,
follower_active=False, live_modes=live, wifi_notice=notice,
reload_pending=reload)
# (running, current_mode, inputs, now) -> HELD or (source, mode)
MID_TABLE = [
# Nothing changed.
(CLOCK, "clock", _mid(), 500.0, HELD),
(CLOCK, "clock", _mid(live=()), 500.0, HELD),
# A game went live: it takes the panel (the first live mode).
(CLOCK, "clock", _mid(live=("nfl_live", "nhl_live")), 500.0, (LIVE, "nfl_live")),
# ... even with a notice pending: the claim is made now, and the next
# pass shows the notice first (top-of-pass order), then the game.
(CLOCK, "clock", _mid(live=("nfl_live",), notice=FRESH), 500.0, (LIVE, "nfl_live")),
# ... but not over the schedule or an on-demand session.
(CLOCK, "clock", _mid(live=("nfl_live",), schedule_on=False), 500.0, (OFF, None)),
(CLOCK, "x", _mid(live=("nfl_live",), on_demand=True), 500.0, (ONDEM, "x")),
# A screen already on a live mode is not taken over by another.
(CLOCK, "nfl_live", _mid(live=("nfl_live", "nhl_live")), 500.0, (ROTATION, "nfl_live")),
(NFL, "nfl_live", _mid(live=("nhl_live",)), 500.0, HELD),
# The mode moved under the screen: on-demand started, or ended, or the
# rotation was rebuilt.
(CLOCK, "x", _mid(on_demand=True), 500.0, (ONDEM, "x")),
(OD, "weather", _mid(), 500.0, (ROTATION, "weather")),
(CLOCK, "weather", _mid(), 500.0, (ROTATION, "weather")),
# An on-demand session that ends on the same mode keeps the screen.
(OD, "x", _mid(), 500.0, HELD),
# The schedule: off ends it; an on-demand override holds.
(CLOCK, "clock", _mid(schedule_on=False), 500.0, (OFF, None)),
(OD, "x", _mid(on_demand=True, schedule_on=False), 500.0, HELD),
# A WiFi notice, compared with its expiry; on-demand outranks it.
(CLOCK, "clock", _mid(notice=FRESH), 999.9, (WIFI, None)),
(CLOCK, "clock", _mid(notice=FRESH), 1_000.0, HELD),
(NFL, "nfl_live", _mid(notice=FRESH), 500.0, (WIFI, None)),
(OD, "x", _mid(on_demand=True, notice=FRESH), 500.0, HELD),
# A plugin reload waits at the top of the loop.
(CLOCK, "clock", _mid(reload=True), 500.0, (Source.RELOAD, None)),
(OD, "x", _mid(on_demand=True, reload=True), 500.0, (Source.RELOAD, None)),
# The order between them: a moved mode before the schedule, the
# schedule before a notice, a notice before a reload.
(CLOCK, "weather", _mid(schedule_on=False), 500.0, (ROTATION, "weather")),
(CLOCK, "clock", _mid(schedule_on=False, notice=FRESH), 500.0, (OFF, None)),
(CLOCK, "clock", _mid(notice=FRESH, reload=True), 500.0, (WIFI, None)),
]
class TestMidScreen:
@pytest.mark.parametrize("running,current,inputs,now,expected", MID_TABLE)
def test_decide(self, running, current, inputs, now, expected):
state = ArbiterState(current_mode=current)
plan = Arbiter.decide(state, inputs, now, running=running)
if expected == HELD:
assert plan is running
else:
assert plan is not running
assert (plan.source, plan.mode) == expected
def test_the_running_plans(self):
assert (CLOCK.source, CLOCK.mode) == (ROTATION, "clock")
assert (NFL.source, NFL.mode) == (LIVE, "nfl_live")
assert LIVE not in NFL.preemptible_by
assert (OD.source, OD.mode) == (ONDEM, "x")
def test_a_follower_and_vegas_never_preempt(self):
for plan in (CLOCK, NFL, OD):
assert FOLLOW not in plan.preemptible_by
assert LEGACY not in plan.preemptible_by
def test_nothing_in_preemptible_by_means_nothing_preempts(self):
bare = replace(CLOCK, preemptible_by=frozenset())
inputs = _mid(live=("nfl_live",), schedule_on=False, notice=FRESH, reload=True)
assert Arbiter.decide(ArbiterState(current_mode="weather"), inputs, 0.0,
running=bare) is bare
def test_mid_screen_decide_reads_no_clock(self):
boom = MagicMock(side_effect=AssertionError("decide read the clock"))
with patch("time.time", boom), patch("time.monotonic", boom):
for running, current, inputs, now, _ in MID_TABLE:
Arbiter.decide(ArbiterState(current_mode=current), inputs, now,
running=running)
# (current, inputs) -> live_takeover. The mode a mid-screen check claims.
TAKEOVER_TABLE = [
("clock", _mid(live=("nfl_live",)), "nfl_live"),
("clock", _mid(live=()), None),
("clock", _mid(live=None), None), # no scan was due
("nfl_live", _mid(live=("nfl_live",)), None),
("clock", _mid(live=("nfl_live",), on_demand=True), None),
("clock", _mid(live=("nfl_live",), schedule_on=False), None),
("clock", replace(_mid(live=("nfl_live",)), vegas_enabled=True,
vegas_live_in_ticker=True), None),
]
@pytest.mark.parametrize("current,inputs,expected", TAKEOVER_TABLE)
def test_live_takeover(current, inputs, expected):
assert live_takeover(ArbiterState(current_mode=current), inputs) == expected
assert step("07:00:30") is LEGACY
-142
View File
@@ -1,142 +0,0 @@
"""The report of a scrolling screen held by its plugin's update().
While a plugin's update() runs it holds the plugin's lock, and its screen's
frames are skipped -- on a scroller, a frozen strip -- with nothing logged.
_note_display_hold times each such run and reports one of
DISPLAY_HOLD_REPORT_SECONDS or more.
"""
import threading
import types
from unittest.mock import MagicMock
import pytest
class _Clock:
"""display_controller's clock: moves only when run() sleeps or a test says."""
def __init__(self, start=10_000.0):
self.t = start
def now(self):
return self.t
def sleep(self, seconds):
self.t += max(seconds, 0.0005)
def module(self):
return types.SimpleNamespace(time=self.now, monotonic=self.now,
perf_counter=self.now, sleep=self.sleep)
@pytest.fixture
def clock(monkeypatch):
c = _Clock()
monkeypatch.setattr("src.display_controller.time", c.module())
return c
class _Locks:
"""get_plugin_lock for one plugin, whose lock the test can hold."""
def __init__(self):
self.lock = threading.Lock()
def __call__(self, plugin_id):
return self.lock
@pytest.fixture
def held(test_display_controller):
c = test_display_controller
locks = _Locks()
c.plugin_manager.get_plugin_lock = locks
c.plugin_manager._warn_rate_limited = MagicMock()
c.plugin_manager.health_tracker = MagicMock()
c._display_hold = None
return c, locks.lock
def _plugin(plugin_id):
p = MagicMock()
p.plugin_id = plugin_id
p.display.return_value = True
return p
class TestTheDisplayHoldReport:
def test_a_long_hold_is_reported_when_it_ends(self, held, clock):
c, lock = held
ticker = _plugin("ticker")
lock.acquire() # update() running
assert c._display_once(ticker, "ticker", False, report_hold=True) is True
clock.t += 0.2
assert c._display_once(ticker, "ticker", False, report_hold=True) is True
assert ticker.display.call_count == 0
c.plugin_manager._warn_rate_limited.assert_not_called()
clock.t += 0.2
lock.release() # update() done
c._display_once(ticker, "ticker", False, report_hold=True)
assert ticker.display.call_count == 1
key, message, plugin_id, ms = c.plugin_manager._warn_rate_limited.call_args[0]
assert key == "display-hold:ticker" and plugin_id == "ticker"
assert "held" in message and ms == pytest.approx(400.0)
c.plugin_manager.health_tracker.record_busy_skip.assert_called_once_with(
"ticker", "display hold", pytest.approx(0.4))
def test_a_short_hold_is_not(self, held, clock):
c, lock = held
ticker = _plugin("ticker")
lock.acquire()
c._display_once(ticker, "ticker", False, report_hold=True)
clock.t += 0.1
lock.release()
c._display_once(ticker, "ticker", False, report_hold=True)
c.plugin_manager._warn_rate_limited.assert_not_called()
c.plugin_manager.health_tracker.record_busy_skip.assert_not_called()
def test_a_hold_that_ends_on_another_plugins_screen_is_not_blamed_on_it(
self, held, clock):
c, lock = held
lock.acquire()
c._display_once(_plugin("ticker"), "ticker", False, report_hold=True)
clock.t += 1.0
lock.release()
c._display_once(_plugin("clock"), "clock", False, report_hold=True)
c.plugin_manager._warn_rate_limited.assert_not_called()
assert c._display_hold is None
def test_frames_that_draw_report_nothing(self, held, clock):
c, _lock = held
ticker = _plugin("ticker")
for _ in range(5):
c._display_once(ticker, "ticker", False, report_hold=True)
clock.t += 0.5
c.plugin_manager._warn_rate_limited.assert_not_called()
def test_the_1hz_loop_reports_no_holds(self, held, clock):
# A static screen's frames are a second apart: one skipped frame is
# not a measured hold, and nothing on the panel froze. (On ledpi the
# first version reported every such skip as "held 1000 ms".)
c, lock = held
board = _plugin("board")
lock.acquire()
c._display_once(board, "board", False)
clock.t += 1.0
lock.release()
c._display_once(board, "board", False)
c.plugin_manager._warn_rate_limited.assert_not_called()
c.plugin_manager.health_tracker.record_busy_skip.assert_not_called()
def test_a_run_left_open_is_dropped_by_a_1hz_frame(self, held, clock):
c, lock = held
ticker = _plugin("ticker")
lock.acquire()
c._display_once(ticker, "ticker", False, report_hold=True)
lock.release()
c._display_once(ticker, "ticker", False) # the 1 Hz loop draws
assert c._display_hold is None
clock.t += 5.0
c._display_once(ticker, "ticker", False, report_hold=True)
c.plugin_manager._warn_rate_limited.assert_not_called()
-151
View File
@@ -430,157 +430,6 @@ class TestClear:
assert "clear request" in response.get_json()["message"]
CLIENT = "web_interface.blueprints.api_v3.control_client"
def _mailbox_file(shared_cache):
_, _, directory = shared_cache
return directory / f"{ERROR_CLEAR_REQUEST_KEY}.json"
class TestClearOverTheSocket:
"""``errors.clear``: the display applies the clear before it answers, and
the mailbox is written only when the socket could not carry it."""
@pytest.fixture
def socket_up(self, display, monkeypatch):
"""The control socket, as the display serves it: errors_clear runs
the display's own handler against its publisher."""
from src.ipc import client as control_client
from src.ipc.contract import ErrorsClearArgs
_, publisher, _ = display
monkeypatch.setattr(errors, "_snapshot_publisher", publisher)
calls = []
def errors_clear(request_id, cutoff, **kw):
calls.append((request_id, cutoff))
return errors.apply_error_clear(request_id, ErrorsClearArgs(cutoff=cutoff))
monkeypatch.setattr(f"{CLIENT}.errors_clear", errors_clear)
assert control_client.errors_clear is errors_clear
return calls
def test_socket_clear_is_applied_before_the_answer(self, web, display, socket_up,
shared_cache):
aggregator, publisher, _ = display
for _ in range(3):
_fail(aggregator)
publisher.tick()
response = web.post("/api/v3/errors/clear", json={"all": True})
assert response.status_code == 200
body = response.get_json()
data = body["data"]
assert data["transport"] == "socket" and data["applied"] is True
assert data["cleared_count"] == 3
assert body["message"] == "Cleared all errors"
[(request_id, _)] = socket_up
assert data["request_id"] == request_id
# Applied already: no display tick needed, nothing pending.
assert aggregator.get_error_summary()["total_errors"] == 0
summary = _summary(web)
assert summary["total_errors"] == 0 and summary["clear_pending"] is False
# And no mailbox file.
assert not _mailbox_file(shared_cache).exists()
def test_an_older_mailbox_request_does_not_read_as_pending(self, web, display, socket_up,
shared_cache, monkeypatch):
# A clear that went to the mailbox while the socket was down, then a
# wider one over the socket: the old request has nothing left to hide.
from unittest.mock import patch
from src.ipc import client as control_client
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
with patch(f"{CLIENT}.errors_clear",
side_effect=control_client.ControlError("no_socket")):
data = web.post("/api/v3/errors/clear", json={"max_age_hours": 1}).get_json()["data"]
assert data["transport"] == "mailbox"
assert _mailbox_file(shared_cache).exists()
data = web.post("/api/v3/errors/clear", json={"all": True}).get_json()["data"]
assert data["transport"] == "socket"
assert _summary(web)["clear_pending"] is False
@pytest.mark.parametrize("reason", ["no_socket", "refused", "disabled", "unsupported"])
def test_no_socket_writes_the_mailbox(self, web, display, shared_cache, monkeypatch, reason):
from src.ipc import client as control_client
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(
side_effect=control_client.ControlError(reason, sent=False)))
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
data = web.post("/api/v3/errors/clear", json={"all": True}).get_json()["data"]
assert data["transport"] == "mailbox" and data["applied"] is False
assert _mailbox_file(shared_cache).exists()
assert _summary(web)["clear_pending"] is True
publisher.tick()
assert aggregator.get_error_summary()["total_errors"] == 0
def test_an_older_display_gets_the_mailbox(self, web, display, shared_cache, monkeypatch):
# The upgrade case: new web interface, a display from before errors.clear.
from src.ipc import client as control_client
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(
side_effect=control_client.ControlError("unknown_command", sent=True)))
aggregator, publisher, _ = display
_fail(aggregator)
publisher.tick()
data = web.post("/api/v3/errors/clear", json={"all": True}).get_json()["data"]
assert data["transport"] == "mailbox"
publisher.tick()
assert aggregator.get_error_summary()["total_errors"] == 0
@pytest.mark.parametrize("reason", ["internal", "timeout", "busy", "invalid_args"])
def test_a_display_that_had_it_and_failed_is_an_error(self, web, display, shared_cache,
monkeypatch, reason):
from src.ipc import client as control_client
monkeypatch.setattr(f"{CLIENT}.errors_clear", MagicMock(
side_effect=control_client.ControlError(reason, sent=True)))
response = web.post("/api/v3/errors/clear", json={"all": True})
assert response.status_code == 503
assert response.get_json()["context"]["socket_error"] == reason
assert not _mailbox_file(shared_cache).exists()
class TestPublisherMailboxPoll:
def test_the_mailbox_is_read_only_when_its_file_changed(self, display, shared_cache):
_, publisher, _ = display
_, web_cache, _ = shared_cache
publisher.tick()
publisher.cache_manager = MagicMock(wraps=publisher.cache_manager)
def reads():
return [c for c in publisher.cache_manager.get.call_args_list
if c.args[0] == ERROR_CLEAR_REQUEST_KEY]
for _ in range(5):
publisher.tick()
assert reads() == [] # no file: a stat per tick, no read
errors.request_error_clear(web_cache, 1.0)
publisher.tick()
assert len(reads()) == 1
for _ in range(5):
publisher.tick()
assert len(reads()) == 1 # unchanged file: not read again
errors.request_error_clear(web_cache, 2.0)
publisher.tick()
assert len(reads()) == 2
def test_clear_now_publishes_what_it_applied(self, display, shared_cache):
aggregator, publisher, _ = display
_, web_cache, _ = shared_cache
_fail(aggregator)
assert publisher.clear_now("sock-1", datetime.now().timestamp() + 1) == 1
snapshot = web_cache.get(ERROR_SNAPSHOT_KEY, max_age=None, memory_ttl=0)
assert snapshot["applied_clear_id"] == "sock-1"
assert snapshot["total_errors"] == 0
assert snapshot["applied_clear_cutoff"] is not None
def test_the_handler_needs_a_running_publisher(self, monkeypatch):
from src.ipc.contract import ErrorsClearArgs
monkeypatch.setattr(errors, "_snapshot_publisher", None)
with pytest.raises(RuntimeError):
errors.apply_error_clear("x", ErrorsClearArgs(cutoff=1.0))
@pytest.mark.skipif(not hasattr(os, "fchmod") or os.name == "nt",
reason="POSIX file modes")
def test_both_files_are_group_readable(web, display, shared_cache):
-36
View File
@@ -576,42 +576,6 @@ class TestCounters:
assert snap["hosts"]["site.api.espn.com"]["requests"] == 1
assert snap["totals"]["bytes"] == 3 * len(b'{"ok": 1}')
def test_wire_bytes_are_the_compressed_size(self, service):
# Built the way requests builds a real response: a urllib3
# HTTPResponse carrying a gzip body, decoded when .content is read.
import gzip
import io
from requests.adapters import HTTPAdapter
from urllib3.response import HTTPResponse
decoded = json.dumps({"events": [{"id": str(i), "name": "x" * 200}
for i in range(50)]}).encode()
wire = gzip.compress(decoded)
def handler(url, kwargs):
raw = HTTPResponse(body=io.BytesIO(wire), status=200,
headers={"Content-Encoding": "gzip",
"Content-Type": "application/json"},
preload_content=False, decode_content=True)
request = requests.Request("GET", url).prepare()
response = HTTPAdapter().build_response(request, raw)
response.content # what Session.get does for a non-streamed call
return response
response = service.get(FakeSession(handler), "https://site.api.espn.com/x")
assert response.content == decoded
totals = _counters(service)
assert totals["bytes"] == len(decoded)
assert totals["wire_bytes"] == len(wire) < len(decoded)
def test_wire_bytes_fall_back_to_the_decoded_size(self, service):
# No urllib3 response behind it (a test double, another adapter):
# count what is known rather than nothing.
service.get(FakeSession(), "https://api.test/x")
totals = _counters(service)
assert totals["wire_bytes"] == totals["bytes"] == len(b'{"ok": 1}')
def test_errors_and_http_errors(self, service):
def handler(url, kwargs):
if url.endswith("/down"):
+51
View File
@@ -807,3 +807,54 @@ def test_a_process_with_the_gc_monitor_exits_cleanly():
assert proc.returncode == 0, proc.stderr
assert "Exception ignored" not in proc.stderr
assert "installed at exit: False" in proc.stdout
SLOW = 1 / 110.0 # a panel that cannot reach a 120 Hz cap
def _windows(recorder, n, interval, start=0.0):
for i in range(n):
_feed(recorder, [interval] * 200, start=start + 50.0 * i)
_aggregate(recorder)
def _shortfall_warnings(caplog):
return [r for r in caplog.records
if r.name == "src.common.frame_timing" and "Limit Refresh Rate" in r.getMessage()]
def test_a_panel_slower_than_its_cap_is_reported_once(tmp_path, caplog):
r = _recorder(tmp_path)
r.plan_refresh(120.0)
caplog.set_level("WARNING")
_windows(r, 3, SLOW) # adopted on the 2nd window, checked on the 4th
assert _shortfall_warnings(caplog) == []
_windows(r, 3, SLOW, start=1000.0)
warnings = _shortfall_warnings(caplog)
assert len(warnings) == 1
assert "about 110 Hz" in warnings[0].getMessage()
assert "to 100 Hz" in warnings[0].getMessage()
def test_a_panel_that_reaches_its_cap_is_not_reported(tmp_path, caplog):
r = _recorder(tmp_path)
r.plan_refresh(100.0)
caplog.set_level("WARNING")
_windows(r, 6, PERIOD)
assert _shortfall_warnings(caplog) == []
def test_without_a_planned_rate_nothing_is_checked(tmp_path, caplog):
# The emulator and the fallback canvas: DisplayManager never calls
# plan_refresh(), since their frames are not paced by a panel.
r = _recorder(tmp_path)
caplog.set_level("WARNING")
_windows(r, 6, SLOW)
assert _shortfall_warnings(caplog) == []
def test_the_snapshot_records_the_planned_rate(tmp_path):
r = _recorder(tmp_path)
assert r.snapshot()["planned_refresh_hz"] is None
r.plan_refresh(120.0)
assert r.snapshot()["planned_refresh_hz"] == 120.0
+1 -1
View File
@@ -709,6 +709,6 @@ class TestTheControllersOwnScreens:
c._check_wifi_status_message.return_value = None
inputs = DisplayController._arbiter_inputs(c)
assert inputs.wifi_notice is None
assert Arbiter.decide(ArbiterState(), inputs, 0.0).source is Source.ROTATION
assert Arbiter.decide(ArbiterState(), inputs, 0.0).source is Source.LEGACY
assert dm.is_currently_scrolling()
assert dm._frame_hold == 2
@@ -1,188 +0,0 @@
"""GET /api/v3/plugins/installed never waits on the network for registry data.
The route used `get_registry_info`, which goes through `fetch_registry`: on a
cold (or expired) cache that downloads plugins.json from GitHub with a 10s
timeout and three attempts -- and, with no registry to fall back on, every
plugin's lookup repeated the whole cycle. The first plugin-list load after a
restart waited on GitHub, and offline it waited out every timeout.
Now the route reads the registry copy already in memory, however old, and a
missing or expired copy only starts a background refresh. These tests block
the network at the socket layer (DNS lookups hang, then fail) and assert the
request returns quickly without a single network attempt on the request path.
"""
import functools
import socket
import threading
import time
from unittest.mock import MagicMock
import pytest
import requests
from test._api_v3_test_helpers import ( # noqa: F401 - fixtures
api_v3_client, api_v3_module,
)
from src.plugin_system.store_manager import PluginStoreManager
# How long a blocked lookup hangs before failing: long enough that a single
# one on the request path blows the response budget below.
HANG_SECONDS = 1.5
FAST_SECONDS = 1.0
REGISTRY = {'plugins': [
{'id': 'weather', 'name': 'Weather', 'verified': True, 'latest_version': '1.2.0'},
]}
@pytest.fixture
def blocked_network(monkeypatch):
"""Every DNS lookup hangs, then fails; records the thread it came from."""
attempts = []
def hang_then_fail(host, *args, **kwargs):
attempts.append((host, threading.current_thread().name))
time.sleep(HANG_SECONDS)
raise socket.gaierror(-3, 'Temporary failure in name resolution (blocked by test)')
monkeypatch.setattr(socket, 'getaddrinfo', hang_then_fail)
return attempts
@pytest.fixture
def store(tmp_path, monkeypatch):
store = PluginStoreManager(plugins_dir=str(tmp_path / 'plugins'))
# One attempt, no pause between attempts, so a background refresh against
# the blocked network ends within the test (teardown joins it).
monkeypatch.setattr(store, '_http_get_with_retries', functools.partial(
PluginStoreManager._http_get_with_retries, store, max_retries=1))
yield store
thread = getattr(store, '_registry_refresh_thread', None)
if thread is not None:
thread.join(timeout=30)
@pytest.fixture
def get_installed(api_v3_module, api_v3_client, store, tmp_path):
api = api_v3_module.api_v3
api.plugin_store_manager = store
api.plugin_catalog.plugins_dir = str(tmp_path / 'plugins')
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[
{'id': 'weather', 'name': 'Weather', 'version': '1.0.0'},
{'id': 'clock', 'name': 'Clock', 'version': '2.0.0'},
])
api.plugin_catalog.get_plugin_display_modes = MagicMock(return_value=[])
api.config_manager.load_config = MagicMock(return_value={})
def _get():
start = time.perf_counter()
response = api_v3_client.get('/api/v3/plugins/installed')
elapsed = time.perf_counter() - start
assert response.status_code == 200
plugins = {p['id']: p for p in response.get_json()['data']['plugins']}
return plugins, elapsed
return _get
def _request_path_attempts(attempts):
return [a for a in attempts if a[1] != 'registry-refresh']
def test_a_cold_cache_offline_returns_fast_without_registry_info(get_installed, blocked_network):
plugins, elapsed = get_installed()
assert _request_path_attempts(blocked_network) == []
assert elapsed < FAST_SECONDS, f"installed list took {elapsed:.2f}s with the network blocked"
weather = plugins['weather']
assert weather['latest_version'] == ''
assert weather['update_available'] is False
assert weather['verified'] is False
def test_a_stale_cache_is_used_as_is_without_a_fetch(get_installed, store, blocked_network):
store.registry_cache = REGISTRY
store.registry_cache_time = time.time() - store.registry_cache_timeout - 3600
plugins, elapsed = get_installed()
assert _request_path_attempts(blocked_network) == []
assert elapsed < FAST_SECONDS
weather = plugins['weather']
assert weather['latest_version'] == '1.2.0'
assert weather['update_available'] is True
assert weather['verified'] is True
assert plugins['clock']['latest_version'] == ''
def test_a_fresh_cache_starts_no_refresh(get_installed, store, blocked_network):
store.registry_cache = REGISTRY
store.registry_cache_time = time.time()
plugins, _ = get_installed()
assert blocked_network == []
assert store._registry_refresh_thread is None
assert plugins['weather']['update_available'] is True
def test_a_cold_cache_is_filled_in_the_background_for_the_next_load(get_installed, store, monkeypatch):
response = MagicMock()
response.json.return_value = REGISTRY
fetched_on = []
def fake_get(url, **kwargs):
fetched_on.append(threading.current_thread().name)
return response
monkeypatch.setattr(store, '_http_get_with_retries', fake_get)
first, _ = get_installed()
assert first['weather']['update_available'] is False
store._registry_refresh_thread.join(timeout=10)
second, _ = get_installed()
assert second['weather']['latest_version'] == '1.2.0'
assert second['weather']['update_available'] is True
# One background fetch for the whole listing, none on the request path.
assert fetched_on == ['registry-refresh']
def test_an_offline_background_refresh_backs_off(store, monkeypatch):
def offline(url, **kwargs):
raise requests.ConnectionError('blocked by test')
monkeypatch.setattr(store, '_http_get_with_retries', offline)
assert store.refresh_registry_in_background() is True
store._registry_refresh_thread.join(timeout=10)
assert store.registry_cache is None
# Offline: the next page load does not start another attempt straight away.
assert store.refresh_registry_in_background() is False
store._registry_refresh_retry_after = 0.0
assert store.refresh_registry_in_background() is True
def test_only_one_background_refresh_runs_at_a_time(store, monkeypatch):
release = threading.Event()
def slow(url, **kwargs):
release.wait(10)
raise requests.ConnectionError('blocked by test')
monkeypatch.setattr(store, '_http_get_with_retries', slow)
try:
assert store.refresh_registry_in_background() is True
assert store.refresh_registry_in_background() is False
finally:
release.set()
def test_get_registry_info_still_fetches_for_the_store(store, monkeypatch):
"""The store, install and update paths keep fetching a cold registry."""
response = MagicMock()
response.json.return_value = REGISTRY
monkeypatch.setattr(store, '_http_get_with_retries', MagicMock(return_value=response))
assert store.get_registry_info('weather')['latest_version'] == '1.2.0'
store._http_get_with_retries.assert_called_once()
+1 -2
View File
@@ -167,8 +167,7 @@ class TestOnDemandArgs:
def test_every_command_has_an_argument_type(self, cmd):
args = {Command.ON_DEMAND_START: {'plugin_id': 'p'},
Command.PLUGIN_RELOAD: {'plugin_id': 'p'},
Command.BRIGHTNESS_SET: {'brightness': 50},
Command.ERRORS_CLEAR: {'cutoff': 1790000000.0}}.get(cmd, {})
Command.BRIGHTNESS_SET: {'brightness': 50}}.get(cmd, {})
c.parse_args(cmd, args)
def test_hello_versions(self):
+5 -15
View File
@@ -41,15 +41,6 @@ def _reload(plugin_id):
return _command(Command.PLUGIN_RELOAD, PluginReloadArgs(plugin_id))
def _screen(mode):
"""A rotation screen of ``mode``, as the ScreenRunner hands it to the
1 Hz loop's frame wait."""
from src.display_arbiter import ArbiterState, rotation_plan
from src.screen_runner import Screen
plan = rotation_plan(ArbiterState(current_mode=mode))
return Screen(plan, plugin=None, accepts_display_mode=False, start=0.0)
@pytest.fixture
def dc(test_display_controller):
c = test_display_controller
@@ -262,8 +253,7 @@ class TestRealTimeWake:
dc.on_demand_active = False
dc._activate_on_demand = MagicMock(
side_effect=lambda request: setattr(dc, 'current_display_mode', 'weather'))
dc._wifi_notice_pending = MagicMock(return_value=False) # the dwell's check
dc._read_wifi_notice = MagicMock(return_value=None) # the frame wait's
dc._wifi_notice_pending = MagicMock(return_value=False)
dc._tick_plugin_updates = MagicMock()
dc._check_live_takeover = MagicMock()
dc.cache_manager.get = MagicMock(return_value=None)
@@ -292,10 +282,10 @@ class TestRealTimeWake:
dc.current_display_mode = 'clock'
stamps = []
t = self._post_later(server, self._start_line(f's{i}'), 0.05, stamps)
ended = dc._wait_frame_interval(1.0, _screen('clock'))
ended = dc._wait_frame_interval(1.0, 'clock')
woke = time.monotonic()
t.join()
assert ended is not None
assert ended is True
latencies.append(woke - stamps[0])
latencies.sort()
print(f"static-screen wake latency: median {latencies[5] * 1000:.2f} ms, "
@@ -325,7 +315,7 @@ class TestRealTimeWake:
'args': {'brightness': 33}}).encode()
started = time.monotonic()
t = self._post_later(server, line, 0.1, stamps)
assert dc._wait_frame_interval(0.5, _screen('clock')) is None
assert dc._wait_frame_interval(0.5, 'clock') is False
assert time.monotonic() - started >= 0.49
t.join()
dc.display_manager.set_brightness.assert_called_once_with(33)
@@ -333,7 +323,7 @@ class TestRealTimeWake:
def test_without_a_socket_it_is_a_plain_sleep(self, dc):
dc._control_server = None
started = time.monotonic()
assert dc._wait_frame_interval(0.2, _screen('clock')) is None
assert dc._wait_frame_interval(0.2, 'clock') is False
assert time.monotonic() - started >= 0.19
-509
View File
@@ -1,509 +0,0 @@
"""Stage 4 of the control socket: the file mailboxes are only a fallback.
* The client knows whether the display had the request (``ControlError.sent``)
and ``should_fall_back`` allows a mailbox write only when it did not, or
when the display is too old to know the command (the upgrade case).
* ``errors.clear`` is answered on the connection thread by a handler the
display registers; a display without one answers like an older display.
* The display looks at the on-demand mailbox once a second while the socket
is up (0.25 s without it), reads it only when its file changed, never
touches it for a socket command, and logs who still writes it.
* ``CacheManager.file_signature`` / ``MailboxWatch`` make a look one stat().
The web routes are covered in test_api_v3_on_demand_socket.py and
test_error_snapshot_cross_process.py.
"""
import json
import logging
import os
import socket
import time
from unittest.mock import MagicMock, patch
import pytest
from src.cache_manager import CacheManager, MailboxWatch
from src.ipc import client
from src.ipc import contract as c
from src.ipc.contract import Command, ErrorsClearArgs, OnDemandStartArgs, ProtocolError
from src.ipc.server import ControlServer, QueuedCommand
MAILBOX = 'display_on_demand_request'
# -- the client: was the request sent? ---------------------------------------------
class FakeSock:
"""Stands in for a connected socket in client._exchange."""
def __init__(self, replies=(), send_error=None, recv_error=None):
self.replies = list(replies)
self.send_error = send_error
self.recv_error = recv_error
self.sent = b''
def settimeout(self, _t):
pass
def sendall(self, data):
if self.send_error is not None:
raise self.send_error
self.sent += data
def recv(self, _n):
if self.recv_error is not None:
raise self.recv_error
return self.replies.pop(0) if self.replies else b''
def close(self):
pass
def _reply(request_id, **body):
return (json.dumps(dict({'v': 1, 'id': request_id}, **body)) + '\n').encode()
def _call(sock=None, connect_error=None, request_id='rid-1'):
with patch.object(client, 'socket_supported', return_value=True), \
patch.object(client, '_connect',
side_effect=connect_error, return_value=sock):
return client.request(Command.PING, {}, request_id=request_id, paths=['/x.sock'])
def _error(**kw):
with pytest.raises(client.ControlError) as e:
_call(**kw)
return e.value
class TestSent:
def test_an_answer_is_returned(self):
assert _call(FakeSock([_reply('rid-1', ok=True, result={'pong': True})])) == {'pong': True}
@pytest.mark.parametrize('reason', ['no_socket', 'refused', 'timeout', 'busy'])
def test_a_failed_connect_was_not_sent(self, reason):
e = _error(connect_error=client.ControlError(reason))
assert e.reason == reason and e.sent is False
assert client.should_fall_back(e)
def test_a_send_that_timed_out_was_not_sent(self):
e = _error(sock=FakeSock(send_error=socket.timeout()))
assert e.reason == 'timeout' and e.sent is False
assert client.should_fall_back(e)
def test_silence_after_the_request_was_sent(self):
e = _error(sock=FakeSock(recv_error=socket.timeout()))
assert e.reason == 'timeout' and e.sent is True
assert not client.should_fall_back(e)
def test_a_hang_up_after_the_request_was_sent(self):
e = _error(sock=FakeSock([]))
assert e.reason == 'closed' and e.sent is True
assert not client.should_fall_back(e)
def test_a_garbled_reply(self):
e = _error(sock=FakeSock([b'not json\n']))
assert e.reason == 'bad_response' and e.sent is True
assert not client.should_fall_back(e)
@pytest.mark.parametrize('code', ['busy', 'invalid_args', 'internal', 'pending', 'failed'])
def test_a_display_error_with_an_id_was_sent(self, code):
sock = FakeSock([_reply('rid-1', ok=False, error={'code': code, 'message': 'x'})])
e = _error(sock=sock)
assert e.reason == code and e.sent is True
assert not client.should_fall_back(e)
@pytest.mark.parametrize('code', ['forbidden', 'busy'])
def test_a_refusal_at_the_door_was_not_sent(self, code):
# forbidden, or too many connections: answered before the request
# was read, so with no id.
sock = FakeSock([(json.dumps({'v': 1, 'id': None, 'ok': False,
'error': {'code': code, 'message': 'x'}}) + '\n').encode()])
e = _error(sock=sock)
assert e.reason == code and e.sent is False
assert client.should_fall_back(e)
@pytest.mark.parametrize('code', ['unknown_command', 'unsupported_version'])
def test_an_older_display_is_fallen_back_from(self, code):
sock = FakeSock([_reply('rid-1', ok=False, error={'code': code, 'message': 'x'})])
e = _error(sock=sock)
assert e.sent is True
assert client.should_fall_back(e)
def test_a_request_refused_locally_never_left(self):
with pytest.raises(client.ControlError) as e:
client.request(Command.ERRORS_CLEAR, {'cutoff': 'soon'}, paths=['/x.sock'])
assert e.value.reason == 'invalid_request' and e.value.sent is False
assert client.should_fall_back(e.value)
def test_a_client_bug_falls_back(self):
assert client.should_fall_back(RuntimeError('boom'))
# -- errors.clear on the server ------------------------------------------------------
def _line(cmd, args, rid='r1'):
return c.encode_message({'v': 1, 'id': rid, 'cmd': cmd, 'args': args})
class TestErrorsClearOnTheServer:
def test_the_handler_answers_it(self, tmp_path):
seen = []
def handler(request_id, args):
seen.append((request_id, args))
return {'request_id': request_id, 'cutoff': args.cutoff, 'cleared': 4}
server = ControlServer(str(tmp_path / 's.sock'),
handlers={Command.ERRORS_CLEAR: handler})
response = server.handle_line(_line(Command.ERRORS_CLEAR, {'cutoff': 123}))
assert response.ok and response.result == {'request_id': 'r1', 'cutoff': 123.0,
'cleared': 4}
assert seen == [('r1', ErrorsClearArgs(cutoff=123.0))]
assert not server.has_pending # not queued for the render thread
def test_a_display_without_a_handler_answers_like_an_older_one(self, tmp_path):
server = ControlServer(str(tmp_path / 's.sock'))
response = server.handle_line(_line(Command.ERRORS_CLEAR, {'cutoff': 1}))
assert not response.ok and response.error.code == c.ErrorCode.UNKNOWN_COMMAND
def test_only_direct_commands_take_a_handler(self, tmp_path):
server = ControlServer(str(tmp_path / 's.sock'),
handlers={Command.ON_DEMAND_START: lambda *a: {}})
response = server.handle_line(_line(Command.ON_DEMAND_START, {'plugin_id': 'p'}))
assert response.ok and response.result['accepted'] is True # still queued
assert server.has_pending
def test_a_handler_error_is_contained(self, tmp_path):
def boom(*_a):
raise ValueError('disk gone')
server = ControlServer(str(tmp_path / 's.sock'), handlers={Command.ERRORS_CLEAR: boom})
response = server.handle_line(_line(Command.ERRORS_CLEAR, {'cutoff': 1}))
assert response.error.code == c.ErrorCode.INTERNAL
assert 'disk gone' not in response.error.message
def test_a_handler_can_refuse_with_a_code(self, tmp_path):
def refuse(*_a):
raise ProtocolError(c.ErrorCode.BUSY, 'later')
server = ControlServer(str(tmp_path / 's.sock'), handlers={Command.ERRORS_CLEAR: refuse})
assert server.handle_line(
_line(Command.ERRORS_CLEAR, {'cutoff': 1})).error.code == c.ErrorCode.BUSY
@pytest.mark.parametrize('cutoff', ['1', None, True, float('inf'), -1])
def test_bad_cutoffs_are_refused(self, cutoff):
with pytest.raises(ProtocolError) as e:
ErrorsClearArgs.from_dict({'cutoff': cutoff})
assert e.value.code == c.ErrorCode.INVALID_ARGS
def test_hello_lists_it(self, tmp_path):
server = ControlServer(str(tmp_path / 's.sock'))
result = server.handle_line(_line(Command.HELLO, {'versions': [1]})).result
assert Command.ERRORS_CLEAR in result['commands']
# -- the display's mailbox poll ------------------------------------------------------
class SignedCache:
"""The slice of CacheManager the poll uses, counting what it costs."""
def __init__(self):
self.data = {}
self.writes = 0
self.version = {}
self.reads = []
self.stats = 0
self.deletes = []
self.sets = []
def file_signature(self, key):
self.stats += 1
return (self.version[key], 0, 0) if key in self.data else None
def get(self, key, *a, **kw):
self.reads.append(key)
return self.data.get(key)
def set(self, key, value, *a, **kw):
self.sets.append(key)
self.data[key] = value
self.writes += 1
self.version[key] = self.writes
def delete(self, key):
self.deletes.append(key)
self.data.pop(key, None)
class FakeServer:
def __init__(self):
self.commands = []
@property
def has_pending(self):
return bool(self.commands)
def drain(self):
out, self.commands = self.commands, []
return out
class Clock:
def __init__(self):
self.t = 1000.0
def __call__(self):
return self.t
@pytest.fixture
def controller(test_display_controller, monkeypatch):
dc = test_display_controller
dc.cache_manager = SignedCache()
dc._activate_on_demand = MagicMock()
dc.on_demand_active = False
dc.on_demand_request_id = None
dc._last_on_demand_poll = None
dc._on_demand_mailbox = None
dc._mailbox_writers_logged = frozenset()
clock = Clock()
monkeypatch.setattr('src.display_controller.time.monotonic', clock)
dc.clock = clock
return dc
def _post(dc, rid, action='start', **fields):
dc.cache_manager.set(MAILBOX, dict({'request_id': rid, 'action': action}, **fields))
def _poll_for(dc, seconds, step=1 / 16): # exact in binary: no drift past a floor
end = dc.clock.t + seconds
while dc.clock.t < end:
dc._poll_on_demand_requests()
dc.clock.t += step
class TestMailboxCadence:
def test_without_a_socket_it_is_looked_at_every_quarter_second(self, controller):
controller._control_server = None
_poll_for(controller, 10.0)
assert 38 <= controller.cache_manager.stats <= 42
def test_with_a_socket_it_is_looked_at_once_a_second(self, controller):
controller._control_server = FakeServer()
_poll_for(controller, 10.0)
assert 9 <= controller.cache_manager.stats <= 11
def test_a_look_that_finds_nothing_reads_nothing(self, controller):
controller._control_server = FakeServer()
_poll_for(controller, 10.0)
assert controller.cache_manager.reads == []
def test_an_unchanged_mailbox_is_not_read_again(self, controller):
# An already-processed start the delete could not remove, say.
controller._control_server = FakeServer()
controller.cache_manager.delete = MagicMock() # the file stays
_post(controller, 'once', plugin_id='clock')
_poll_for(controller, 10.0)
assert controller.cache_manager.reads.count(MAILBOX) <= 2 # the read + the re-check
controller._activate_on_demand.assert_called_once()
def test_a_mailbox_request_lands_within_a_second_with_the_socket_up(self, controller):
# The upgrade case the other way round: a new display, and a web
# interface (or a plugin) that still writes the mailbox.
controller._control_server = FakeServer()
controller._poll_on_demand_requests()
controller.clock.t += 0.1
_post(controller, 'old-web', plugin_id='clock')
posted = controller.clock.t
while not controller._activate_on_demand.called:
controller._poll_on_demand_requests()
controller.clock.t += 0.05
assert controller.clock.t - posted < 1.5
assert controller.clock.t - posted <= controller.MAILBOX_POLL_INTERVAL_WITH_SOCKET + 0.06
assert MAILBOX in controller.cache_manager.deletes # consumed
def test_socket_commands_still_land_at_once(self, controller):
server = controller._control_server = FakeServer()
controller._poll_on_demand_requests()
server.commands.append(QueuedCommand('sock', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests() # inside the mailbox interval
controller._activate_on_demand.assert_called_once()
class TestSocketCommandsLeaveTheMailboxAlone:
def test_a_socket_start_reads_and_deletes_no_mailbox(self, controller):
server = controller._control_server = FakeServer()
controller._poll_on_demand_requests()
before = list(controller.cache_manager.reads)
server.commands.append(QueuedCommand('s1', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests()
controller._activate_on_demand.assert_called_once()
assert MAILBOX not in controller.cache_manager.reads[len(before):]
assert controller.cache_manager.deletes == []
def test_a_socket_stop_reads_and_deletes_no_mailbox(self, controller):
from src.ipc.contract import OnDemandStopArgs
controller.on_demand_active = True
controller._clear_on_demand = MagicMock()
server = controller._control_server = FakeServer()
server.commands.append(QueuedCommand('s2', Command.ON_DEMAND_STOP,
OnDemandStopArgs(), time.time()))
controller.clock.t += 5
controller._drain_control_commands()
controller._clear_on_demand.assert_called_once()
assert MAILBOX not in controller.cache_manager.reads
assert controller.cache_manager.deletes == []
def test_a_mailbox_copy_of_a_socket_command_is_dropped(self, controller):
# An older web interface timed out after the display queued the
# command, then wrote the mailbox too.
server = controller._control_server = FakeServer()
server.commands.append(QueuedCommand('both', Command.ON_DEMAND_START,
OnDemandStartArgs(plugin_id='clock'), time.time()))
controller._poll_on_demand_requests()
_post(controller, 'both', plugin_id='clock')
_poll_for(controller, 2.0)
controller._activate_on_demand.assert_called_once()
assert MAILBOX not in controller.cache_manager.data
class TestDeprecationLog:
def test_each_mailbox_writer_is_logged_once(self, controller, caplog):
controller._control_server = FakeServer()
caplog.set_level(logging.INFO, logger='src.display_controller')
for i, plugin in enumerate(['on-air', 'on-air', 'pomodoro-timer']):
_post(controller, f'r{i}', plugin_id=plugin)
_poll_for(controller, 1.2)
lines = [r.getMessage() for r in caplog.records if 'file mailbox' in r.getMessage()]
assert len(lines) == 2
assert 'on-air' in lines[0] and 'pomodoro-timer' in lines[1]
def test_nothing_is_logged_without_a_socket(self, controller, caplog):
controller._control_server = None
caplog.set_level(logging.INFO, logger='src.display_controller')
_post(controller, 'r', plugin_id='on-air')
_poll_for(controller, 1.0)
controller._activate_on_demand.assert_called_once()
assert not [r for r in caplog.records if 'file mailbox' in r.getMessage()]
# -- file_signature and MailboxWatch -------------------------------------------------
@pytest.fixture
def real_cache(tmp_path, monkeypatch):
monkeypatch.setattr(CacheManager, '_get_writable_cache_dir', lambda self: str(tmp_path))
cache = CacheManager()
yield cache
cache.stop_cleanup_thread()
class TestFileSignature:
def test_absent_key(self, real_cache):
assert real_cache.file_signature('nothing') is None
def test_every_write_is_a_new_signature(self, real_cache):
seen = set()
for i in range(20):
# Same size each time, written as fast as possible.
real_cache.set(MAILBOX, {'request_id': f'r{i:02d}'})
sig = real_cache.file_signature(MAILBOX)
assert isinstance(sig, tuple)
seen.add(sig)
assert len(seen) == 20
def test_gone_after_a_delete(self, real_cache):
real_cache.set(MAILBOX, {'a': 1})
real_cache.delete(MAILBOX)
assert real_cache.file_signature(MAILBOX) is None
class TestMailboxWatch:
def test_reads_once_per_write(self, real_cache):
watch = MailboxWatch(MAILBOX)
assert watch.changed(real_cache) is False # no file
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache) is True
assert watch.changed(real_cache) is False
real_cache.set(MAILBOX, {'request_id': 'b'})
assert watch.changed(real_cache) is True
def test_forget_reads_again(self, real_cache):
watch = MailboxWatch(MAILBOX)
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache) is True
watch.forget()
assert watch.changed(real_cache) is True
def test_a_rewrite_after_a_delete_is_seen(self, real_cache):
watch = MailboxWatch(MAILBOX)
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache)
real_cache.delete(MAILBOX)
assert watch.changed(real_cache) is False
real_cache.set(MAILBOX, {'request_id': 'a'})
assert watch.changed(real_cache) is True
def test_a_cache_that_cannot_tell_is_read_every_time(self):
watch = MailboxWatch(MAILBOX)
assert watch.changed(MagicMock()) is True
assert watch.changed(MagicMock()) is True
assert watch.changed(object()) is True
# -- end to end over a real socket ---------------------------------------------------
@pytest.mark.skipif(not c.socket_supported(), reason='AF_UNIX sockets are Linux/macOS only')
class TestOverTheSocket:
@pytest.fixture
def sock_path(self):
import shutil
import tempfile
d = tempfile.mkdtemp(prefix='lmipc-')
yield os.path.join(d, 'control.sock')
shutil.rmtree(d, ignore_errors=True)
def test_errors_clear_round_trip(self, sock_path):
def handler(request_id, args):
return {'request_id': request_id, 'cutoff': args.cutoff, 'cleared': 2}
server = ControlServer(sock_path, handlers={Command.ERRORS_CLEAR: handler})
assert server.start()
try:
result = client.errors_clear('clr-1', 1790000000.0, paths=[sock_path])
assert result == {'request_id': 'clr-1', 'cutoff': 1790000000.0, 'cleared': 2}
finally:
server.close()
def test_an_older_display_is_an_upgrade_fallback(self, sock_path):
server = ControlServer(sock_path) # no errors.clear handler
assert server.start()
try:
with pytest.raises(client.ControlError) as e:
client.errors_clear('clr-2', 1.0, paths=[sock_path])
assert e.value.reason == 'unknown_command' and e.value.sent is True
assert client.should_fall_back(e.value)
finally:
server.close()
def test_no_display_is_a_fallback(self, sock_path):
with pytest.raises(client.ControlError) as e:
client.errors_clear('clr-3', 1.0, paths=[sock_path])
assert e.value.reason == 'no_socket' and e.value.sent is False
assert client.should_fall_back(e.value)
def test_a_full_queue_is_not_a_fallback(self, sock_path):
server = ControlServer(sock_path, queue_size=1)
assert server.start()
try:
client.on_demand_start('q1', 'clock', None, paths=[sock_path])
with pytest.raises(client.ControlError) as e:
client.on_demand_start('q2', 'clock', None, paths=[sock_path])
assert e.value.reason == 'busy' and e.value.sent is True
assert not client.should_fall_back(e.value)
finally:
server.close()
+2 -4
View File
@@ -87,10 +87,8 @@ class TestANamedLiveModeIsShown:
def test_the_named_mode_survives_a_restart(self, football):
football._activate_on_demand({'plugin_id': 'football-scoreboard',
'mode': 'ncaa_fb_live'})
# The last on-demand config write, not the last write of any key: the
# font-usage publisher thread writes its own key at its own pace.
saved = [c for c in football.cache_manager.set.call_args_list
if c.args and c.args[0] == 'display_on_demand_config'][-1]
saved = football.cache_manager.set.call_args_list[-1]
assert saved.args[0] == 'display_on_demand_config'
config = saved.args[1]
assert config['named_mode'] == 'ncaa_fb_live'
+1 -1
View File
@@ -423,7 +423,7 @@ def web_listing(api_v3_module, api_v3_client, shared_cache, tmp_path): # noqa:
{"id": "clock", "name": "Clock", "version": "1.1.0"},
{"id": "weather", "name": "Weather", "version": "3.0.0"},
])
api.plugin_store_manager.get_cached_registry_info = MagicMock(return_value=None)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.plugin_store_manager._get_local_git_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={
"clock": {"enabled": True}, "weather": {"enabled": True}})
-455
View File
@@ -1,455 +0,0 @@
"""ScreenRunner (src/screen_runner.py) on a scripted host and a fake clock.
The golden traces (test_run_loop_golden.py) run it inside the real
DisplayController; these pin down its own contract: which ExitReason each
way of ending gives, the order it calls its host in, how the 125 Hz loop
paces, and which service point asks what.
"""
from typing import Any, List, Optional, Tuple
import pytest
from src.display_arbiter import (
RELOAD_PLAN, SCREEN_PREEMPTERS, FramePolicy, ScreenPlan, Source,
)
from src.screen_runner import (
AFTER_COMPLETED_LOOP, AFTER_LOOP, DYNAMIC_GRACE, FINAL, FRAME, HIGH_FPS_INTERVAL,
Checkpoint, ExitReason, FirstFrame, NoticeRead, Screen, ScreenRunner,
)
WIFI_PLAN = ScreenPlan(Source.WIFI)
class Clock:
"""time/perf_counter read ``now``; sleep advances it."""
def __init__(self):
self.now = 0.0
self.sleeps: List[float] = []
def time(self) -> float:
return self.now
def perf_counter(self) -> float:
return self.now
def sleep(self, seconds: float) -> None:
self.sleeps.append(round(seconds, 6))
self.now += seconds
class Host:
"""A ScreenHost whose answers are scripted; records every call."""
def __init__(self, clock: Clock, *, shown=True, raised=False, minimum=10.0,
maximum=10.0, dynamic=False, policy=FramePolicy.STATIC,
completed=True, draws=None, checks=None, work=0.0,
cycle_complete_at=None, dwell_to=None):
self.clock = clock
self.calls: List[Tuple[Any, ...]] = []
self.shown, self.raised = shown, raised
self.minimum, self.maximum = minimum, maximum
self.dynamic, self.policy, self.completed = dynamic, policy, completed
#: display() results for the frames after the first, in order.
self.draws = list(draws or [])
#: (checkpoint name, frame number) -> plan, or a callable(screen, cp).
self.checks = checks or {}
self.work = work
self.cycle_complete_at = cycle_complete_at
self.dwell_to = dwell_to
self.frames = 0
def first_frame(self, plan, plugin):
self.calls.append(("first", plan.mode))
return FirstFrame(self.shown, self.raised, False)
def complete_plan(self, plan, plugin):
self.calls.append(("complete",))
if not self.completed:
return None
return ScreenPlan(plan.source, mode=plan.mode, min_duration=self.minimum,
max_duration=self.maximum, dynamic=self.dynamic,
frame_policy=self.policy, preemptible_by=plan.preemptible_by)
def draw(self, screen):
self.frames += 1
self.calls.append(("draw", self.frames))
self.clock.now += self.work
return self.draws.pop(0) if self.draws else True
def after_frame(self, screen):
self.calls.append(("after_frame",))
def tick(self):
self.calls.append(("tick",))
def service(self, screen):
self.calls.append(("service",))
return ("scan", self.frames)
def wait_frame(self, interval, screen):
self.calls.append(("wait", interval))
self.clock.sleep(interval)
return self._scripted("wait", screen)
def check(self, screen, checkpoint, live_scan=None):
self.calls.append(("check", checkpoint.name, live_scan))
return self._scripted(checkpoint.name, screen, checkpoint)
def _scripted(self, name, screen, checkpoint=None):
answer = self.checks.get((name, self.frames))
if callable(answer):
return answer(screen, checkpoint)
return answer
def dwell(self, seconds):
self.calls.append(("dwell", round(seconds, 6)))
self.clock.now = self.dwell_to if self.dwell_to is not None else self.clock.now + seconds
def cycle_complete(self, screen):
self.calls.append(("cycle?",))
return self.cycle_complete_at is not None and self.clock.now >= self.cycle_complete_at
PLAN = ScreenPlan(Source.ROTATION, mode="clock", preemptible_by=SCREEN_PREEMPTERS)
def _run(**kwargs) -> Tuple[Any, Host, Clock]:
clock = Clock()
host = Host(clock, **kwargs)
outcome = ScreenRunner(clock, host).run(PLAN, plugin=object())
return outcome, host, clock
def _names(host, *kinds):
return [c for c in host.calls if c[0] in kinds]
class TestFirstFrame:
def test_no_plugin_is_empty_without_a_dispatch(self):
clock = Clock()
host = Host(clock)
outcome = ScreenRunner(clock, host).run(PLAN, plugin=None)
assert outcome.exit_reason is ExitReason.EMPTY
assert host.calls == []
@pytest.mark.parametrize("raised,reason", [(False, ExitReason.EMPTY),
(True, ExitReason.ERROR)])
def test_nothing_shown(self, raised, reason):
outcome, host, _ = _run(shown=False, raised=raised)
assert outcome.exit_reason is reason
assert host.calls == [("first", "clock")]
def test_an_on_demand_session_with_no_time_left(self):
outcome, host, _ = _run(completed=False)
assert outcome.exit_reason is ExitReason.PREEMPTED
assert outcome.preempted_by is None
assert host.calls == [("first", "clock"), ("complete",)]
class TestStaticLoop:
def test_runs_its_duration_one_frame_a_second(self):
outcome, host, clock = _run(maximum=5.0)
assert outcome.exit_reason is ExitReason.DURATION
assert outcome.elapsed == 5.0
# Frames at 1..4 s; the wait that reaches 5 s ends it before a fifth.
assert host.frames == 4
assert clock.sleeps == [1.0] * 5
# Each frame: wait, tick, draw, follower frame, service, the check.
assert host.calls[2:9] == [("wait", 1.0), ("tick",), ("draw", 1), ("after_frame",),
("service",), ("check", "frame", ("scan", 1)), ("wait", 1.0)]
# A completed loop: the after-loop look reads no notice, then FINAL.
assert host.calls[-2:] == [("check", "after-completed-loop", None), ("check", "final", None)]
def test_preempted_between_frames(self):
by = ScreenPlan(Source.ON_DEMAND, mode="x")
outcome, host, _ = _run(maximum=30.0, checks={("frame", 3): by})
assert outcome.exit_reason is ExitReason.PREEMPTED
assert outcome.preempted_by is by
assert outcome.elapsed == 3.0
# Decided at the service point: no second look, no dwell.
assert host.calls[-1] == ("check", "frame", ("scan", 3))
def test_preempted_by_a_socket_command_in_the_frame_wait(self):
outcome, host, _ = _run(maximum=30.0, checks={("wait", 2): WIFI_PLAN})
assert outcome.exit_reason is ExitReason.PREEMPTED
assert host.calls[-1] == ("wait", 1.0)
def test_display_false_makes_up_the_minimum(self):
outcome, host, _ = _run(maximum=12.0, draws=[True, False])
assert outcome.exit_reason is ExitReason.DISPLAY_FALSE
# The loop ended early: the after-loop look may read a notice, then
# the dwell makes up the rest of the 12 s.
assert _names(host, "check", "dwell")[-4:] == [
("check", "after-loop", None), ("dwell", 10.0),
("check", "after-dwell", None), ("check", "final", None)]
assert outcome.elapsed == 12.0
def test_a_notice_that_cuts_the_dwell_short_ends_the_screen(self):
def notice(screen, checkpoint):
assert checkpoint.notice is NoticeRead.ALWAYS
return WIFI_PLAN if checkpoint.notice_counts else None
outcome, _, _ = _run(maximum=12.0, draws=[False], dwell_to=6.0,
checks={("after-dwell", 1): notice})
assert outcome.exit_reason is ExitReason.PREEMPTED
def test_a_notice_after_a_full_dwell_does_not(self):
def notice(screen, checkpoint):
return WIFI_PLAN if checkpoint.notice_counts else None
outcome, _, _ = _run(maximum=12.0, draws=[False],
checks={("after-dwell", 1): notice})
assert outcome.exit_reason is ExitReason.DISPLAY_FALSE
def test_a_reload_ends_the_loop_but_the_screen_counts(self):
outcome, host, _ = _run(maximum=30.0, checks={("frame", 2): RELOAD_PLAN})
assert outcome.exit_reason is ExitReason.RELOAD
# Not decided at the service point: the after-loop look, the dwell
# (which returns at once while a reload waits) and FINAL follow.
assert ("check", "after-loop", None) in host.calls
assert ("check", "final", None) in host.calls
def test_a_change_after_the_loop_preempts(self):
by = ScreenPlan(Source.ROTATION, mode="weather")
outcome, _, _ = _run(maximum=3.0, checks={("final", 2): by})
assert outcome.exit_reason is ExitReason.PREEMPTED
assert outcome.preempted_by is by
def test_a_dynamic_screen_keeps_going_on_false(self):
outcome, host, _ = _run(minimum=3.0, maximum=8.0, dynamic=True,
draws=[False, False, False])
assert outcome.exit_reason is ExitReason.DURATION
assert host.frames == 7
class TestHighFpsLoop:
def test_paces_to_the_deadline(self):
outcome, host, clock = _run(maximum=0.05, policy=FramePolicy.HIGH_FPS, work=0.003)
assert outcome.exit_reason is ExitReason.DURATION
# 3 ms of drawing leaves 5 ms of an 8 ms frame to sleep.
assert set(clock.sleeps) == {round(HIGH_FPS_INTERVAL - 0.003, 6)}
def test_an_overrun_frame_still_yields(self):
_, _, clock = _run(maximum=0.05, policy=FramePolicy.HIGH_FPS, work=0.02)
assert set(clock.sleeps) == {0.001}
def test_service_before_the_sleep_check_after(self):
_, host, clock = _run(maximum=0.016, policy=FramePolicy.HIGH_FPS)
first = host.calls[2:7]
assert first == [("draw", 1), ("after_frame",), ("tick",), ("service",),
("check", "frame", ("scan", 1))]
assert clock.sleeps[0] == HIGH_FPS_INTERVAL
def test_preempted_after_the_frames_sleep(self):
outcome, _, clock = _run(maximum=30.0, policy=FramePolicy.HIGH_FPS,
checks={("frame", 3): WIFI_PLAN})
assert outcome.exit_reason is ExitReason.PREEMPTED
assert clock.now == pytest.approx(3 * HIGH_FPS_INTERVAL)
def test_display_false_has_no_make_up_dwell(self):
outcome, host, _ = _run(maximum=30.0, policy=FramePolicy.HIGH_FPS,
draws=[True, False])
assert outcome.exit_reason is ExitReason.DISPLAY_FALSE
assert not _names(host, "dwell")
class TestDynamicDuration:
def test_cycle_complete_after_the_minimum_and_grace(self):
outcome, host, _ = _run(minimum=3.0, maximum=20.0, dynamic=True,
cycle_complete_at=1.0)
assert outcome.exit_reason is ExitReason.CYCLE_COMPLETE
# Not asked before minimum + grace (3.5 s): the 1 Hz loop's first
# frame at or past it is the one at 4 s.
assert outcome.elapsed == 4.0
assert DYNAMIC_GRACE == 0.5
def test_capped_at_the_maximum(self):
outcome, host, _ = _run(minimum=3.0, maximum=6.0, dynamic=True)
assert outcome.exit_reason is ExitReason.DURATION
# The end-of-screen log asks the plugin once more, before FINAL.
assert host.calls[-3:] == [("check", "after-completed-loop", None), ("cycle?",),
("check", "final", None)]
def test_high_fps_cycle_complete(self):
outcome, _, _ = _run(minimum=0.02, maximum=1.0, dynamic=True,
policy=FramePolicy.HIGH_FPS, cycle_complete_at=0.0)
assert outcome.exit_reason is ExitReason.CYCLE_COMPLETE
assert outcome.elapsed >= 0.02 + DYNAMIC_GRACE
class TestCheckpoints:
@pytest.mark.parametrize("checkpoint,notice,reload", [
(FRAME, NoticeRead.IF_UNDECIDED, True),
(AFTER_LOOP, NoticeRead.IF_UNDECIDED, False),
(AFTER_COMPLETED_LOOP, NoticeRead.NEVER, False),
(FINAL, NoticeRead.NEVER, False),
])
def test_what_each_service_point_considers(self, checkpoint, notice, reload):
"""A reload counts only between frames; the notice file is read
where the loop read it before stage 3."""
assert isinstance(checkpoint, Checkpoint)
assert (checkpoint.notice, checkpoint.reload) == (notice, reload)
def test_the_screen_carries_the_completed_plan(self):
seen: List[Optional[Screen]] = []
def grab(screen, checkpoint):
seen.append(screen)
_run(maximum=2.0, checks={("frame", 1): grab})
assert seen[0].plan.frame_policy is FramePolicy.STATIC
assert seen[0].mode == "clock"
class TestControllerServicePoint:
"""DisplayController._screen_check: the reads it makes and what it claims,
on a controller built by the run-loop harness."""
@pytest.fixture
def dc(self, tmp_path):
import os
os.environ.setdefault("EMULATOR", "true")
from test._run_loop_harness import FakePlugin, RunLoopHarness
h = RunLoopHarness(tmp_path, horizon=10)
h.add_plugin(FakePlugin("clock", ["clock"], duration=20))
h.add_plugin(FakePlugin("sports", ["sports_live"], duration=20,
live=(0, 100), live_priority=True))
dc = h.controller
dc.current_display_mode = "clock"
dc.current_mode_index = 0
dc.reads = []
def read():
dc.reads.append(1)
return {"message": "AP mode", "expires_at": 1e12}
dc._check_wifi_status_message = read
return dc
@staticmethod
def _screen(mode="clock"):
from src.display_arbiter import ArbiterState, rotation_plan
return Screen(rotation_plan(ArbiterState(current_mode=mode)), plugin=None,
accepts_display_mode=False, start=0.0)
def test_a_pending_notice_is_read_and_ends_the_screen(self, dc):
by = dc._screen_check(self._screen(), FRAME)
assert by.source is Source.WIFI and len(dc.reads) == 1
def test_not_read_once_the_mode_has_moved(self, dc):
dc.current_display_mode = "sports_live"
by = dc._screen_check(self._screen(), FRAME)
assert by.source is Source.ROTATION and dc.reads == []
def test_not_read_while_scheduled_off(self, dc):
dc.is_display_active = False
by = dc._screen_check(self._screen(), FRAME)
assert by.source is Source.SCHEDULED_OFF and dc.reads == []
def test_not_read_during_on_demand(self, dc):
dc.on_demand_active = True
dc.on_demand_schedule_override = True
assert dc._screen_check(self._screen(), FRAME) is None
assert dc.reads == []
def test_a_live_takeover_is_claimed_before_the_notice(self, dc):
by = dc._screen_check(self._screen(), FRAME, live_scan=("sports_live",))
assert by.source is Source.LIVE and dc.reads == []
assert dc.current_display_mode == "sports_live"
assert dc._live_takeover_unshown is True and dc._live_resume_index == 0
def test_after_a_completed_loop_the_notice_is_not_read(self, dc):
assert dc._screen_check(self._screen(), AFTER_COMPLETED_LOOP) is None
assert dc.reads == []
def test_after_a_full_dwell_it_is_read_but_does_not_count(self, dc):
from src.screen_runner import after_dwell
assert dc._screen_check(self._screen(), after_dwell(False)) is None
assert len(dc.reads) == 1
def test_a_reload_counts_only_between_frames(self, dc):
dc._check_wifi_status_message = lambda: None
dc._pending_plugin_reloads = ("pending",)
assert dc._screen_check(self._screen(), FRAME) is RELOAD_PLAN
assert dc._screen_check(self._screen(), AFTER_LOOP) is None
def test_an_on_demand_index_past_a_shortened_list_starts_again(self, dc):
"""_take_plan writes the shown index back, so the session's next
step goes on from the mode actually shown."""
from src.display_arbiter import Arbiter, ArbiterInputs
dc.on_demand_active = True
dc.on_demand_modes = ["clock", "sports_live"]
dc.on_demand_mode_index = 5 # the list shrank under it
inputs = ArbiterInputs(schedule_on=True, on_demand_active=True,
follower_active=False)
plan = dc._take_plan(Arbiter.decide(dc._arbiter_state(), inputs, 0.0))
assert plan.mode == "clock" and dc.on_demand_mode_index == 0
dc._advance_on_demand()
assert dc.current_display_mode == "sports_live"
@pytest.mark.parametrize("policy,report_hold", [(FramePolicy.HIGH_FPS, True),
(FramePolicy.STATIC, False)])
def test_only_the_high_fps_loop_reports_a_held_frame(policy, report_hold):
"""#758's report_hold: the 125 Hz loop times frames held by update();
the 1 Hz loop's frames are a second apart and must not."""
from unittest.mock import MagicMock
from src.display_controller import _ScreenHost
controller = MagicMock()
plan = ScreenPlan(Source.ROTATION, mode="ticker", frame_policy=policy)
plugin = object()
_ScreenHost(controller).draw(Screen(plan, plugin, True, 0.0))
controller._display_once.assert_called_once_with(plugin, "ticker", True,
report_hold=report_hold)
def _rows(tmp_path, horizon, build):
import os
os.environ.setdefault("EMULATOR", "true")
from test._run_loop_harness import RunLoopHarness
h = RunLoopHarness(tmp_path, horizon=horizon)
build(h)
return h.run()["screens"]
class TestThroughRun:
"""Service points that only the full loop reaches, on the harness."""
def test_a_notice_pending_when_a_later_frame_is_empty_ends_the_screen(self, tmp_path):
"""The after-loop look reads the notice when the loop ended early: a
1 Hz screen whose second frame has nothing to show must not sit in
its make-up dwell (which only notices a notice that arrives during
it) while the notice expires."""
from test._run_loop_harness import FakePlugin
def build(h):
h.add_plugin(FakePlugin("flaky", ["flaky"], duration=12, first_frame_only=True))
h.add_plugin(FakePlugin("clock", ["clock"], duration=12))
h.wifi_message(0.5, "Connected to HomeNet", duration=5)
rows = _rows(tmp_path, 30, build)
assert rows[0][1] == "flaky" and rows[0][2] == 1.0
assert rows[1][1] == "<wifi>"
# ... and the mode it cut short comes back, not the next one.
after = next(row for row in rows[1:] if row[1] != "<wifi>")
assert after[1] == "flaky"
def test_vegas_yielding_to_a_follower_shows_a_rotation_screen_first(self, tmp_path):
"""Pins today's behaviour (docs/RUN_LOOP_REDESIGN.md, "may be
wrong"): the interrupt check stops the ticker for a follower, but
the yield path never looks at one, so a rotation screen runs its
full duration before the next pass hands the panel to the leader."""
from test._run_loop_harness import FakePlugin
def build(h):
h.add_plugin(FakePlugin("clock", ["clock"], duration=20))
h.add_plugin(FakePlugin("weather", ["weather"], duration=20))
h.enable_vegas(cycle=30)
h.sync.follower_windows = [(10, 60)]
rows = _rows(tmp_path, 50, build)
assert rows[0][1] == "<vegas>" and rows[0][3] == "vegas-interrupt"
assert 10.0 <= rows[0][0] + rows[0][2] <= 10.2
assert rows[1][1] == "clock" and rows[1][2] == 20.0
assert rows[2][1] == "<follower>"
+44
View File
@@ -21,6 +21,7 @@ from src.common.scroll_config import ( # noqa: E402
refresh_hz_from_config,
resolve,
)
from src.common import scroll_config # noqa: E402
class FakeHelper:
@@ -504,3 +505,46 @@ class TestSpeedAdvice:
got = solve_crisp(50, 125.74)
assert got.steppiness == "smooth"
assert got.pixels_per_frame == 1
class TestRefreshShortfall:
"""A panel that cannot reach its cap runs every scroll slow."""
def test_the_ledmatrix_rig_is_told_to_cap_at_100(self):
# Pi 4, 2x128x64 on adafruit-hat-pwm under a 120 Hz cap: measured
# 107.6-113.1 Hz, and frame_timing reports the fast end.
s = scroll_config.refresh_shortfall(113.1, 120)
assert s == {"measured_hz": 113.1, "planned_hz": 120.0,
"suggested_cap_hz": 100, "slow_percent": 6}
def test_a_panel_that_holds_its_cap_is_fine(self):
assert scroll_config.refresh_shortfall(99.95, 100) is None
assert scroll_config.refresh_shortfall(97.5, 100) is None
def test_a_panel_that_beats_its_cap_is_fine(self):
assert scroll_config.refresh_shortfall(125.7, 120) is None
def test_nothing_measured_says_nothing(self):
assert scroll_config.refresh_shortfall(None, 120) is None
assert scroll_config.refresh_shortfall(0, 120) is None
assert scroll_config.refresh_shortfall("fast", 120) is None
def test_the_suggestion_leaves_headroom_under_the_measurement(self):
assert scroll_config.holdable_cap(113.1) == 100
assert scroll_config.holdable_cap(95.0) == 90
# 5% under 105 is 99.75: 100 would sit inside the panel's drift.
assert scroll_config.holdable_cap(105.0) == 90
assert scroll_config.holdable_cap(9.0) is None
assert scroll_config.holdable_cap(None) is None
def test_the_log_line_names_the_cap_to_use(self):
text = scroll_config.describe_refresh_shortfall(
scroll_config.refresh_shortfall(113.1, 120))
assert "about 113 Hz" in text and "120 Hz" in text
assert "6% slower" in text
assert "Set Limit Refresh Rate to 100 Hz" in text
def test_no_suggestion_for_a_panel_too_slow_for_any_cap(self):
text = scroll_config.describe_refresh_shortfall(
scroll_config.refresh_shortfall(9.0, 100))
assert "Set Limit Refresh Rate" not in text
+1 -1
View File
@@ -446,7 +446,7 @@ class TestInstalledPluginsApi:
**(manifest_extra or {})}
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[info])
api.plugin_catalog.get_plugin_directory = MagicMock(return_value=None)
api.plugin_store_manager.get_cached_registry_info = MagicMock(return_value=None)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={'demo': config})
response = api_v3_client.get('/api/v3/plugins/installed')
assert response.status_code == 200
+3 -3
View File
@@ -551,7 +551,7 @@ class TestPluginsAPI:
mock_plugin_catalog.get_all_plugin_info.return_value = [
{'id': 'weather', 'name': 'Weather Plugin'}
]
api_v3.plugin_store_manager.get_cached_registry_info.return_value = None
api_v3.plugin_store_manager.get_registry_info.return_value = None
response = client.get('/api/v3/plugins/installed')
@@ -571,7 +571,7 @@ class TestPluginsAPI:
{'id': 'weather', 'name': 'Weather', 'version': '1.0.0'}
]
# Registry advertises a newer version than the installed one.
api_v3.plugin_store_manager.get_cached_registry_info.return_value = {
api_v3.plugin_store_manager.get_registry_info.return_value = {
'verified': True, 'latest_version': '1.2.0'
}
@@ -592,7 +592,7 @@ class TestPluginsAPI:
mock_plugin_catalog.get_all_plugin_info.return_value = [
{'id': 'weather', 'name': 'Weather', 'version': '1.2.0'}
]
api_v3.plugin_store_manager.get_cached_registry_info.return_value = {
api_v3.plugin_store_manager.get_registry_info.return_value = {
'verified': True, 'latest_version': '1.2.0'
}
+1 -1
View File
@@ -59,7 +59,7 @@ class TestInstalledList:
info = {"id": "demo", "name": "Demo", "version": "1.0.0",
"description": "stale cached copy", "loaded": False}
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[info])
api.plugin_store_manager.get_cached_registry_info = MagicMock(return_value=None)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.plugin_store_manager._get_local_git_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={})
@@ -0,0 +1,63 @@
"""GET /api/v3/config/refresh-rate: the cap, the measured rate, a cap to hold."""
import json
from unittest.mock import MagicMock
import pytest
from flask import Flask
from web_interface.blueprints.api_v3 import api_v3
@pytest.fixture
def client(monkeypatch, tmp_path):
stats = tmp_path / "stats.json"
monkeypatch.setattr("src.common.frame_timing.default_stats_path", lambda: str(stats))
manager = MagicMock()
manager.load_config.return_value = {
"display": {"hardware": {"limit_refresh_rate_hz": 120}}}
monkeypatch.setattr(api_v3, "config_manager", manager, raising=False)
app = Flask(__name__)
app.register_blueprint(api_v3, url_prefix="/api/v3")
c = app.test_client()
c.stats_path = stats
return c
def _get(client):
body = client.get("/api/v3/config/refresh-rate").get_json()
assert body["status"] == "success"
return body["data"]
def test_nothing_measured_yet(client):
data = _get(client)
assert data == {"planned_hz": 120.0, "measured_hz": None, "shortfall": None}
def test_a_panel_short_of_its_cap_gets_a_cap_it_can_hold(client):
client.stats_path.write_text(json.dumps(
{"measured_refresh_hz": 110.4, "planned_refresh_hz": 120.0}))
data = _get(client)
assert data["measured_hz"] == 110.4
assert data["shortfall"]["suggested_cap_hz"] == 100
assert data["shortfall"]["slow_percent"] == 8
def test_a_panel_at_its_cap_has_no_shortfall(client):
client.stats_path.write_text(json.dumps(
{"measured_refresh_hz": 121.3, "planned_refresh_hz": 120.0}))
assert _get(client)["shortfall"] is None
def test_a_file_written_under_another_cap_is_stale(client):
# The cap was changed to 120 but the display still runs under 100 Hz.
client.stats_path.write_text(json.dumps(
{"measured_refresh_hz": 99.9, "planned_refresh_hz": 100.0}))
data = _get(client)
assert data["measured_hz"] is None
assert data["shortfall"] is None
def test_a_file_from_a_display_too_old_to_record_its_cap_still_counts(client):
client.stats_path.write_text(json.dumps({"measured_refresh_hz": 110.4}))
assert _get(client)["shortfall"]["suggested_cap_hz"] == 100
@@ -26,44 +26,13 @@ import pytest
@pytest.fixture
def client(monkeypatch):
from web_interface import app as web_app
def client():
from web_interface.app import app
# The captive-portal before_request hook shells out to systemctl/nmcli
# whenever its 30s cache is cold, so on a Linux host whether a request
# here runs subprocess depended on how long ago the previous one was --
# and several tests below stub subprocess. Pin it: no test in this file
# is about AP mode.
monkeypatch.setattr(web_app, 'is_ap_mode_active', lambda: False)
# GET /plugins/installed looks up each plugin's registry entry, and on a
# cold cache that fetches plugins.json from GitHub -- so without a
# connection those tests sat in the HTTP retry loop. No test in this
# file is about the registry.
monkeypatch.setattr(web_app.plugin_store_manager, 'fetch_registry',
lambda *args, **kwargs: {'plugins': []})
app.config['TESTING'] = True
with app.test_client() as c:
yield c
@pytest.fixture(autouse=True)
def starlark_apps_dir(tmp_path, monkeypatch):
"""Point every Starlark storage path at tmp_path for every test.
The manifest, its directory and the lock file are three separate module
constants. Fixtures that redirected the first two but not the lock left
the lock pointing at the repo, and on Linux (where the lock is taken)
each test created starlark-apps/manifest.json.lock in the checkout. The
directory is not created here; tests that need it make it.
"""
from web_interface.blueprints import api_v3 as module
apps_dir = tmp_path / "starlark-apps"
monkeypatch.setattr(module, '_STARLARK_APPS_DIR', apps_dir)
monkeypatch.setattr(module, '_STARLARK_MANIFEST_FILE', apps_dir / 'manifest.json')
monkeypatch.setattr(module, '_STARLARK_MANIFEST_LOCK_FILE', apps_dir / 'manifest.json.lock')
return apps_dir
class TestRoutesAreRegistered:
"""The failure was a missing route, so check the URL map directly.
@@ -422,9 +391,13 @@ class TestTheManifestSurvivesConcurrentWriters:
"""
@pytest.fixture
def starlark_dir(self, starlark_apps_dir):
starlark_apps_dir.mkdir()
return starlark_apps_dir
def starlark_dir(self, tmp_path, monkeypatch):
from web_interface.blueprints import api_v3 as module
apps_dir = tmp_path / "starlark-apps"
apps_dir.mkdir()
monkeypatch.setattr(module, '_STARLARK_APPS_DIR', apps_dir)
monkeypatch.setattr(module, '_STARLARK_MANIFEST_FILE', apps_dir / 'manifest.json')
return apps_dir
def test_each_writer_gets_its_own_temp_file(self, starlark_dir):
from web_interface.blueprints import api_v3 as module
@@ -580,9 +553,13 @@ class TestTheManifestStaysRelocatable:
"""
@pytest.fixture
def starlark_dir(self, starlark_apps_dir):
starlark_apps_dir.mkdir()
return starlark_apps_dir
def starlark_dir(self, tmp_path, monkeypatch):
from web_interface.blueprints import api_v3 as module
apps_dir = tmp_path / "starlark-apps"
apps_dir.mkdir()
monkeypatch.setattr(module, '_STARLARK_APPS_DIR', apps_dir)
monkeypatch.setattr(module, '_STARLARK_MANIFEST_FILE', apps_dir / 'manifest.json')
return apps_dir
def _install(self, tmp_path):
from web_interface.blueprints import api_v3 as module
@@ -622,10 +599,13 @@ class TestManifestLockPreventsLostUpdates:
"""
@pytest.fixture
def starlark_dir(self, starlark_apps_dir):
def starlark_dir(self, tmp_path, monkeypatch):
from web_interface.blueprints import api_v3 as module
apps_dir = starlark_apps_dir
apps_dir = tmp_path / "starlark-apps"
apps_dir.mkdir()
monkeypatch.setattr(module, '_STARLARK_APPS_DIR', apps_dir)
monkeypatch.setattr(module, '_STARLARK_MANIFEST_FILE', apps_dir / 'manifest.json')
monkeypatch.setattr(module, '_STARLARK_MANIFEST_LOCK_FILE', apps_dir / 'manifest.json.lock')
module._write_starlark_manifest({'apps': {}})
return apps_dir
@@ -713,10 +693,12 @@ class TestConfigAndManifestStayInSync:
"""
@pytest.fixture
def app_dir(self, starlark_apps_dir):
def app_dir(self, tmp_path, monkeypatch):
from web_interface.blueprints import api_v3 as module
apps_dir = starlark_apps_dir
apps_dir = tmp_path / "starlark-apps"
apps_dir.mkdir()
monkeypatch.setattr(module, '_STARLARK_APPS_DIR', apps_dir)
monkeypatch.setattr(module, '_STARLARK_MANIFEST_FILE', apps_dir / 'manifest.json')
one_app_dir = apps_dir / 'demo'
one_app_dir.mkdir()
module._write_starlark_manifest({'apps': {'demo': {'name': 'Demo', 'enabled': True}}})
@@ -1060,18 +1042,13 @@ class TestTheStoreUsesTheTokenTheUserConfigured:
/plugins/store/github-status reported `authenticated: true` with a
rate_limit of 5000 while /starlark/repository/browse reported a limit of
60 -- the store going blank was that 60 running out.
The managers are attributes web_interface/app.py hangs on the blueprint
when it is imported, so they exist only once some earlier test has
imported the app. Every patch here passes create=True: these tests must
not depend on which test ran before them.
"""
def test_the_store_managers_token_is_used(self):
from web_interface.blueprints import api_v3 as mod
with patch.object(mod.api_v3, 'plugin_store_manager',
MagicMock(github_token='ghp_configured'), create=True):
MagicMock(github_token='ghp_configured')):
assert mod._starlark_github_token() == 'ghp_configured'
def test_a_hand_edited_config_key_still_works(self):
@@ -1080,8 +1057,8 @@ class TestTheStoreUsesTheTokenTheUserConfigured:
cfg = MagicMock()
cfg.load_config.return_value = {'github_token': 'ghp_by_hand'}
with patch.object(mod.api_v3, 'plugin_store_manager',
MagicMock(github_token=None), create=True), \
patch.object(mod.api_v3, 'config_manager', cfg, create=True):
MagicMock(github_token=None)), \
patch.object(mod.api_v3, 'config_manager', cfg):
assert mod._starlark_github_token() == 'ghp_by_hand'
def test_no_token_anywhere_is_not_an_error(self):
@@ -1090,8 +1067,8 @@ class TestTheStoreUsesTheTokenTheUserConfigured:
cfg = MagicMock()
cfg.load_config.return_value = {}
with patch.object(mod.api_v3, 'plugin_store_manager',
MagicMock(github_token=None), create=True), \
patch.object(mod.api_v3, 'config_manager', cfg, create=True):
MagicMock(github_token=None)), \
patch.object(mod.api_v3, 'config_manager', cfg):
assert mod._starlark_github_token() is None
def test_an_unreadable_config_does_not_take_the_store_down(self):
@@ -1100,8 +1077,8 @@ class TestTheStoreUsesTheTokenTheUserConfigured:
cfg = MagicMock()
cfg.load_config.side_effect = OSError("config.json is unreadable")
with patch.object(mod.api_v3, 'plugin_store_manager',
MagicMock(github_token=None), create=True), \
patch.object(mod.api_v3, 'config_manager', cfg, create=True):
MagicMock(github_token=None)), \
patch.object(mod.api_v3, 'config_manager', cfg):
assert mod._starlark_github_token() is None
def test_browse_hands_the_token_to_the_repository(self, client):
@@ -1115,7 +1092,7 @@ class TestTheStoreUsesTheTokenTheUserConfigured:
repo.return_value.get_rate_limit_info.return_value = {'remaining': 4999}
with patch.object(mod.api_v3, 'plugin_store_manager',
MagicMock(github_token='ghp_configured'), create=True), \
MagicMock(github_token='ghp_configured')), \
patch('web_interface.blueprints.api_v3._get_tronbyte_repository_class',
return_value=repo):
client.get('/api/v3/starlark/repository/browse')
@@ -1159,23 +1136,12 @@ class TestPixletEditorHostDefaultsButDoesNotOverride:
captured['env'] = env
return FakeProcess()
# Swap the route module's own ``subprocess`` binding, not the shared
# ``subprocess.Popen``: patching the attribute on the real module is
# process-wide, and the app's before_request hook (the captive-portal
# check) runs ``subprocess.run`` -- ``with Popen(...)`` -- whenever its
# 30s AP-mode cache is cold on a host with systemctl. On the Linux CI
# runner that handed it this FakeProcess and 500'd the request, but
# only when the previous request was more than 30s earlier.
fake_subprocess = types.ModuleType('subprocess')
fake_subprocess.__dict__.update(mod.subprocess.__dict__)
fake_subprocess.Popen = fake_popen
with patch.object(mod, '_validate_starlark_app_path',
return_value=(app_dir, None)), \
patch.object(mod, '_PIXLET_EDITOR_SCRIPT', script), \
patch.object(mod, '_PIXLET_EDITOR_STATE', state_file), \
patch.object(mod, '_find_pixlet_binary', return_value='/usr/bin/pixlet'), \
patch.object(mod, 'subprocess', fake_subprocess), \
patch.object(mod.subprocess, 'Popen', side_effect=fake_popen), \
patch.dict(os.environ):
if operator_host is None:
os.environ.pop('PIXLET_EDITOR_HOST', None)
@@ -1205,13 +1171,15 @@ class TestStandaloneRenderUsesTheDeviceLocation:
SCHEMA = {"schema": [{"typeOf": "location", "id": "location"}]}
@pytest.fixture
def app_dir(self, starlark_apps_dir, monkeypatch):
def app_dir(self, tmp_path, monkeypatch):
from web_interface.blueprints import api_v3 as module
apps_dir = starlark_apps_dir
apps_dir = tmp_path / "starlark-apps"
app_dir = apps_dir / "weather"
app_dir.mkdir(parents=True)
(app_dir / "weather.star").write_text("# app")
(app_dir / "schema.json").write_text(json.dumps(self.SCHEMA))
monkeypatch.setattr(module, '_STARLARK_APPS_DIR', apps_dir)
monkeypatch.setattr(module, '_STARLARK_MANIFEST_FILE', apps_dir / 'manifest.json')
(apps_dir / 'manifest.json').write_text(json.dumps(
{'apps': {'weather': {'star_file': 'weather.star'}}}))
config_manager = MagicMock()
@@ -128,7 +128,6 @@ class Web:
store = api.plugin_store_manager
store.plugins_dir = str(self.plugins_dir)
store.get_registry_info.return_value = None
store.get_cached_registry_info.return_value = None
store.get_plugin_info.return_value = None
store._get_local_git_info.return_value = None
store.install_plugin.return_value = True
+31 -3
View File
@@ -158,11 +158,16 @@ def _panel_refresh_hz(config):
cap = scroll_config.refresh_hz_from_config(config)
try:
with open(frame_timing.default_stats_path(), encoding='utf-8') as fh:
measured = float(json.load(fh).get('measured_refresh_hz') or 0)
stats = json.load(fh)
measured = float(stats.get('measured_refresh_hz') or 0)
planned = float(stats.get('planned_refresh_hz') or 0)
except (OSError, ValueError, TypeError, AttributeError):
measured = planned = 0.0
# Reject a stale file from a previous hardware config: one written under
# another cap (the display has not restarted since it changed), or, from
# a display too old to record its cap, a measurement far off this one.
if planned and abs(planned - cap) > 0.5:
measured = 0.0
# Reject a stale file from a previous hardware config: a measurement far
# off the cap says the config changed since it was written.
if measured > 0 and 0.5 * cap <= measured <= 1.5 * cap:
return measured, 'measured'
return cap, 'configured'
@@ -189,6 +194,29 @@ def get_scroll_speed_advice():
return jsonify({'status': 'success', 'data': advice})
@api_v3.route('/config/refresh-rate', methods=['GET'])
def get_refresh_rate():
"""The refresh cap, what the panel measured, and a cap it can hold.
Backs the hint under the Display tab's Limit Refresh Rate field. Scroll
speeds are solved against the cap, so a panel that cannot reach it runs
every scroll slow; ``shortfall`` (None when the panel keeps up, or nothing
has been measured yet) says by how much and suggests a cap.
"""
from src.common import scroll_config
if not api_v3.config_manager:
return jsonify({'status': 'error', 'message': 'Config manager not initialized'}), 500
config = api_v3.config_manager.load_config()
planned = scroll_config.refresh_hz_from_config(config)
hz, source = _panel_refresh_hz(config)
measured = hz if source == 'measured' else None
return jsonify({'status': 'success', 'data': {
'planned_hz': planned,
'measured_hz': round(measured, 1) if measured else None,
'shortfall': scroll_config.refresh_shortfall(measured, planned),
}})
@api_v3.route('/config/schedule', methods=['GET'])
def get_schedule_config():
"""Get current schedule configuration"""
+11 -49
View File
@@ -33,28 +33,16 @@ def _cache_manager():
class _NotDelivered(Exception):
"""The display took an on-demand request over the socket and did not
accept it (``busy``, ``invalid_args``, ...) or did not answer in time."""
def __init__(self, reason):
super().__init__(reason)
self.reason = reason
def _deliver_on_demand(payload):
"""Hand an on-demand request to the display: control socket, else mailbox.
The socket (src/ipc) answers with an acknowledgement as soon as the
display has the command queued for its render thread. The file mailbox
is written only when the socket could not carry the request at all
(``control_client.should_fall_back``): no socket (the display is stopped
or predates it), a refused connection, or a display too old to know the
command. The display reads it within its mailbox poll interval.
A display that had the request and refused it, or did not answer in
time, raises :class:`_NotDelivered`: a mailbox copy would be refused
the same way, or hide a stuck display behind a "success".
display has the command queued for its render thread. Any failure -- no
socket (the display is stopped or predates it), a timeout, a refusal --
writes the file mailbox instead, exactly as before the socket existed;
the display reads it within ON_DEMAND_POLL_INTERVAL. Both carry the same
request_id, so a request that reached the display both ways (a reply
that timed out after the command was queued) is still processed once.
Returns ``(transport, socket_error)``: ``'socket'`` and None, or
``'mailbox'`` and the socket failure's reason code.
@@ -69,10 +57,6 @@ def _deliver_on_demand(payload):
return 'socket', None
except control_client.ControlError as e:
reason = _socket_reason_code(e.reason)
if not control_client.should_fall_back(e):
logger.warning("The display did not accept on-demand %s %s (%s)",
payload['action'], payload['request_id'], e)
raise _NotDelivered(reason) from None
if reason in _QUIET_SOCKET_REASONS:
logger.debug("On-demand %s via the mailbox: %s", payload['action'], e)
else:
@@ -85,17 +69,6 @@ def _deliver_on_demand(payload):
return 'mailbox', reason
def _not_delivered_response(request_id, action, reason):
"""The answer when the display had the request and did not accept it."""
status = 400 if reason == 'invalid_args' else 503
return jsonify({
'status': 'error',
'message': (f'The display service did not accept the on-demand {action} '
f'request ({reason})'),
'data': {'request_id': request_id, 'transport': 'socket', 'socket_error': reason},
}), status
def _withdraw_on_demand(request_id):
"""Take a start request the route has refused back out of the mailbox.
@@ -304,10 +277,7 @@ def start_on_demand_display():
'pinned': pinned,
'timestamp': _pkg.time.time()
}
try:
transport, socket_error = _deliver_on_demand(request_payload)
except _NotDelivered as e:
return _not_delivered_response(request_id, 'start', e.reason)
transport, socket_error = _deliver_on_demand(request_payload)
# A socket acknowledgement is the display itself answering: it is
# running and has the request queued, whatever systemd says (a display
@@ -342,10 +312,9 @@ def start_on_demand_display():
# MQTT on-demand command, which posts here with the default -- cold-
# restarted the display process: every plugin reloaded and the panel was
# blank for seconds. The restart bought nothing. The running process
# looks at this mailbox at least once a second, from its dwell sleep,
# its render loops and Vegas's interrupt check as well as the main loop
# (DisplayController._mailbox_poll_interval), and a restarted one got
# the request the same way: the
# reads this mailbox every ON_DEMAND_POLL_INTERVAL (0.25s), from its
# dwell sleep, its render loops and Vegas's interrupt check as well as
# the main loop, and a restarted one got the request the same way: the
# startup path only restores a session the display itself saved
# (display_on_demand_config), so it loaded nothing it would not have had.
service_result = None
@@ -396,14 +365,7 @@ def stop_on_demand_display():
'action': 'stop',
'timestamp': _pkg.time.time()
}
try:
transport, socket_error = _deliver_on_demand(request_payload)
except _NotDelivered as e:
if not stop_service:
return _not_delivered_response(request_id, 'stop', e.reason)
# Stopping the service ends on-demand too, whatever the display did
# with the request.
transport, socket_error = 'socket', e.reason
transport, socket_error = _deliver_on_demand(request_payload)
service_result = None
if stop_service:
+6 -34
View File
@@ -374,23 +374,6 @@ def _read_errors():
return snapshot, clear_request
def _send_error_clear(request_id, cutoff):
"""``errors.clear`` over the control socket: the display's answer, or None
when the socket could not carry it (no socket, or a display older than
the command) and the clear goes to the mailbox instead. A display that
took the request and failed raises ControlError (no second copy)."""
client = _pkg.control_client
try:
return client.errors_clear(request_id, cutoff)
except client.ControlError as e:
if not client.should_fall_back(e):
raise
_pkg._log_socket_failure('errors.clear', e, _pkg._socket_reason_code(e.reason))
except Exception: # never let the socket path break the route
logger.exception("Control socket client failed clearing errors; using the mailbox")
return None
@api_v3.route('/errors/summary', methods=['GET'])
def get_error_summary():
"""
@@ -488,17 +471,7 @@ def clear_old_errors():
now = _pkg.time.time()
cutoff = now if clear_all else now - max_age_hours * 3600
try:
result = _errors.request_error_clear(_errors_cache(), cutoff,
send=_send_error_clear)
except _pkg.control_client.ControlError as e:
reason = _pkg._socket_reason_code(e.reason)
logger.warning("The display did not apply the error clear: %s", e)
return error_response(
error_code=ErrorCode.SYSTEM_ERROR,
message="The display service did not apply the clear",
context={'socket_error': reason},
status_code=503
)
result = _errors.request_error_clear(_errors_cache(), cutoff)
except OSError as e:
logger.error("Could not record an error clear request: %s", e)
return error_response(
@@ -508,12 +481,11 @@ def clear_old_errors():
)
scope = "all errors" if clear_all else f"errors older than {max_age_hours} hours"
if result.get('applied'):
message = f"Cleared {scope}"
else:
message = (f"Clear of {scope} requested; the display service applies it "
f"within about {int(_errors.SNAPSHOT_TICK_INTERVAL)} seconds")
return success_response(data=result, message=message)
return success_response(
data=result,
message=(f"Clear of {scope} requested; the display service applies it "
f"within about {int(_errors.SNAPSHOT_TICK_INTERVAL)} seconds")
)
except Exception as e:
logger.error(f"Error clearing old errors: {e}", exc_info=True)
return error_response(
+2 -7
View File
@@ -106,13 +106,8 @@ def get_installed_plugins():
plugin_config = {}
enabled = bool(plugin_config.get('enabled', False))
# Verified + latest published version from the registry copy already
# in memory. Never a fetch: on a cold cache get_registry_info would
# download plugins.json, and offline wait out its timeout and retries
# for every plugin. With no copy yet these are absent (no update or
# verified badge) and a background refresh fills them in for a
# later load.
store_info = api_v3.plugin_store_manager.get_cached_registry_info(plugin_id)
# Verified + latest published version from registry (no network call)
store_info = api_v3.plugin_store_manager.get_registry_info(plugin_id)
verified = store_info.get('verified', False) if store_info else False
latest_version = store_info.get('latest_version', '') if store_info else ''
installed_version = plugin_info.get('version', '')
@@ -318,6 +318,7 @@
min="0"
max="1000"
class="form-control">
<p id="limit_refresh_rate_hz_hint" class="mt-1 text-xs text-amber-700" aria-live="polite"></p>
</div>
</div>
@@ -915,6 +916,41 @@ document.getElementById('brightness').addEventListener('input', function() {
});
}
// Say so when the panel cannot reach its refresh cap. Scroll speeds are
// worked out against the cap, so every scroll then runs slow, and the
// display has measured a cap the panel can hold.
(function refreshRateHint() {
const hint = document.getElementById('limit_refresh_rate_hz_hint');
const input = document.getElementById('limit_refresh_rate_hz');
if (!hint || !input) return;
fetch('/api/v3/config/refresh-rate')
.then(function(r) { return r.json(); })
.then(function(body) {
const s = body.status === 'success' && body.data.shortfall;
hint.textContent = '';
if (!s) return;
hint.appendChild(document.createTextNode(
'This panel refreshes at about ' + Math.round(s.measured_hz) +
' Hz, below this ' + Math.round(s.planned_hz) + ' Hz cap, so scrolls run about ' +
s.slow_percent + '% slower than set.' + (s.suggested_cap_hz ? ' ' : '')));
if (!s.suggested_cap_hz) return;
const btn = document.createElement('button');
btn.type = 'button';
btn.className = 'underline font-medium';
btn.textContent = 'Use ' + s.suggested_cap_hz + ' Hz';
btn.addEventListener('click', function() {
input.value = s.suggested_cap_hz;
input.dispatchEvent(new Event('input', {bubbles: true}));
input.dispatchEvent(new Event('change', {bubbles: true}));
hint.textContent = 'Save, then restart the display, to apply ' +
s.suggested_cap_hz + ' Hz.';
});
hint.appendChild(btn);
hint.appendChild(document.createTextNode(', a cap it can hold.'));
})
.catch(function() { hint.textContent = ''; });
})();
// Declared before first use: let is not hoisted usably.
let scrollHintTimer = null;
let scrollHintSeq = 0;