feat(ipc): display control socket, stage 1: on-demand with acks

The display process now serves a Unix socket, /run/ledmatrix/control.sock,
carrying versioned newline-delimited JSON commands that are acknowledged.
Stage 1 moves on-demand start/stop (plus status, hello and ping) onto it;
the cache-file mailbox stays as the fallback for one release.

- src/ipc/contract.py: typed request/response envelopes, command args,
  error codes, NDJSON framing with a 64 KiB limit, socket path rules.
- src/ipc/server.py: threaded server owned by the display. Handlers only
  queue onto a bounded queue and ack with the request id; the render
  thread drains it where it reads the mailbox. Bounded clients, timeouts,
  garbage/oversize/disconnect handling; 0660 socket in the cache dir's
  group plus SO_PEERCRED checks; skips cleanly on Windows or when off.
- src/ipc/client.py: one short-timeout request; any failure raises
  ControlError(reason).
- api_v3/display.py: on-demand start/stop try the socket, fall back to
  the mailbox exactly as before, and report transport/socket_error.
- display_controller.py: start/close the server; the mailbox handler body
  is extracted into _handle_on_demand_request and shared by both paths.
- docs/IPC_CONTROL_SOCKET.md: protocol, security model, stage plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-30 21:28:31 -04:00
co-authored by Claude Opus 5.5
parent f4bda50710
commit af8dc3940a
18 changed files with 3129 additions and 42 deletions
+10 -4
View File
@@ -41,7 +41,8 @@ each other. They share three things:
| State | Where | Written by | Read by |
|---|---|---|---|
| On-demand request | cache `display_on_demand_request` | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py) | display: `_poll_on_demand_requests()` |
| On-demand command | control socket `/run/ledmatrix/control.sock` ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py), via [`src/ipc/client.py`](../src/ipc/client.py) | display: [`src/ipc/server.py`](../src/ipc/server.py) acks; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand request (fallback) | cache `display_on_demand_request` | web, when the socket fails; four plugins write it directly | display: `_poll_on_demand_requests()` |
| On-demand state | cache `display_on_demand_state` | display: `_publish_on_demand_state()` | web: `/api/v3/display/on-demand/status` |
| Current screen | cache `display_current_state` | display | web: `/api/v3/display/current-status` |
| Plugin errors | cache `plugin_error_snapshot` | display: `ErrorSnapshotPublisher` ([`src/error_aggregator.py`](../src/error_aggregator.py)) | web: `read_error_report()` for `/api/v3/errors/*` |
@@ -55,9 +56,14 @@ each other. They share three things:
| Render-loop heartbeat | `/run/ledmatrix/display-heartbeat.json` (tmpfs) | display: the render thread, via [`display_watchdog`](../src/display_watchdog.py) | web: `/api/v3/health` (`checks.display_loop`); the update health check |
The on-demand start route starts `ledmatrix.service` when it is not running
(`start_service`, on by default) but never restarts a running one: the display
reads the mailbox every `ON_DEMAND_POLL_INTERVAL` (0.25s), from its dwell
sleep, its render loops and Vegas's interrupt check as well as the main loop.
(`start_service`, on by default) but never restarts a running one. The routes
send the command over the display's control socket and get an ack; when that
fails (a stopped display, one older than the socket) they write the mailbox
instead, which the display reads every `ON_DEMAND_POLL_INTERVAL` (0.25s), from
its dwell sleep, its render loops and Vegas's interrupt check as well as the
main loop. Both ways end in the same handler, `_handle_on_demand_request()`.
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
for the protocol, the permission model and the plan to retire the mailboxes.
### Web and display processes: who runs plugins
+250
View File
@@ -0,0 +1,250 @@
# Control socket (web → display)
The display process serves a Unix socket that the web interface uses to send
it commands and get an answer back. It replaces the cache-file "mailboxes" on
the SD card one command at a time. Stage 1, described here, carries on-demand
start/stop/status. The file mailbox stays as a fallback for one release.
| | |
|---|---|
| Socket | `/run/ledmatrix/control.sock` (tmpfs) |
| Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` |
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop` |
| Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it |
| Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it |
## Why
Before the socket, the web interface sent commands by writing a cache key
(`display_on_demand_request`) that the display read every 0.25 s.
- **No acknowledgement.** The route answered "success" once the file was
written, whether or not a display was running to read it.
- **Lost requests.** The display had to read the request and then delete it.
A request written between those two steps could be thrown away (see
`_consume_on_demand_request`). The cache has no atomic claim to prevent it.
- **Fragile.** Each channel repeated its own permission, atomic-write,
staleness and in-memory-cache rules. Two of them caused bugs: a `memory_ttl`
bug ignored every on-demand request after the first for an hour, and a
stopped display was still reported as "active" for two minutes.
The socket answers every command, carries one request per message (so nothing
can overwrite it), and belongs to the display process. If the display is not
running, the socket does not exist, and the web interface knows right away.
## Protocol (version 1)
**Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at
most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with
`ensure_ascii`, so a newline never appears inside a message. A connection
can carry several requests. Each request gets exactly one response, in order.
**Request**
```json
{"v": 1, "id": "5f0c…", "cmd": "on_demand.start",
"args": {"plugin_id": "clock", "mode": null, "duration": 30, "pinned": false}}
```
- `v` is the protocol version.
- `id` is a printable string of 1-128 characters. It is echoed back in the
response, and for on-demand commands it is also the on-demand `request_id`.
- `cmd` is a command name.
- `args` is an object. It may be omitted when a command takes no arguments.
**Response**
```json
{"v": 1, "id": "5f0c…", "ok": true, "result": {"accepted": true, "request_id": "5f0c…", "queued": 1}}
{"v": 1, "id": "5f0c…", "ok": false, "error": {"code": "busy", "message": "…"}}
```
`id` is `null` only when the request could not be parsed far enough to have
one. Clients branch on `error.code`, never on the message text.
**Commands**
| `cmd` | `args` | `result` | Kind |
|---|---|---|---|
| `hello` | `{versions: [int], client?: str}` | `{version, versions, commands, max_message_bytes, server}` | answered directly |
| `ping` | — | `{pong: true}` | answered directly |
| `on_demand.start` | `{plugin_id?, mode?, duration?, pinned?}` (at least one of `plugin_id` and `mode`) | ack | queued |
| `on_demand.stop` | — | ack | queued |
| `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly |
`duration` is a number of seconds, or a numeric string. `0`, `null` or `""`
mean "until stopped". `pinned` must be a real boolean: the REST route has
already converted strings like `"false"` before it sends the command. The
`on_demand` object in `on_demand.status` is the same dict the display
publishes to `display_on_demand_state`.
**Acknowledgements.** A queued command is *accepted*, not *done*.
`{"accepted": true, "request_id": …}` means the command is waiting in the
render thread's queue, and the render thread will apply it at its next
on-demand check. That is within one frame on a scrolling screen, 0.25 s
during a dwell, and up to 1 s on a static screen, whose frame loop sleeps a
second between frames. Except on a scrolling screen, where the mailbox waits
up to 0.25 s, these are the mailbox's delays too: stage 1 adds
acknowledgements, not speed. Any outcome is published as before
(`display_on_demand_state`, and `status`/`error` for a bad plugin or mode),
and it can be read with `on_demand.status`.
**Versions.** Every request carries `v`. For any command except `hello`, a
`v` the display does not speak gets `unsupported_version`. `hello` is checked
by its `versions` list instead, and its result names the highest version both
sides share, so a client can find out what a display supports before it
relies on anything newer. Stage 1's client sends `v: 1` and falls back to the
mailbox when the display refuses it. It does not send `hello` first, which
saves a round trip.
**Error codes:** `bad_json`, `bad_request`, `message_too_large`,
`unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full,
or too many connections), `forbidden` (peer credentials refused), `internal`.
Try it on a device:
```bash
python3 - <<'EOF'
from src.ipc import client # run from the project directory
print(client.on_demand_status())
EOF
```
## How the display applies a command
The server's threads never touch rendering. A connection thread parses the
request, validates it against the contract, and then does one of two things:
- For a command that changes the panel, it puts a `QueuedCommand` on a
bounded queue (16 entries) and answers with the ack.
- For a query, it answers from a status snapshot the display provides
(`DisplayController._control_status`). The snapshot only reads attributes.
The render thread drains the queue in `_poll_on_demand_requests()`, the same
place it reads the mailbox, and hands each command to
`_handle_on_demand_request()`, which is the mailbox's own handler. The two
paths share all of their code: activation, the processed-id guard, error
publishing, and resuming the rotation afterwards. The 0.25 s floor on the
mailbox read does not apply to the queue, because draining it costs no disk
read. A queued command also lets `_service_pending_changes()` skip its own
floor, so a long scrolling screen or a Vegas iteration takes the command at
its next frame.
**Exactly once.** A command and a mailbox write for the same request share
one `request_id`. If the client times out after the display queued the
command and then also writes the mailbox, the display processes the request
once. The existing `on_demand_request_id` and processed-id checks drop the
second copy.
## Robustness
All of this runs inside the display process, so nothing a client does may
block the render loop or crash it:
- **Bounded connections.** Each connection gets its own daemon thread, with
at most 8 at once. One more is answered `busy` and closed.
- **Timeouts.** Each read and write times out after 2 s. A message must
arrive whole within 5 s of its first byte. An idle connection is closed
after 10 s. A slow or stuck client costs one thread for a few seconds.
- **Malformed input.** A line that is not JSON gets `bad_json`, and the
connection carries on. A line longer than 64 KiB gets `message_too_large`,
and the connection is closed, because the next message boundary cannot be
found. A client that disconnects mid-message is dropped silently. No
exception from a handler leaves the connection thread.
- **Full queue.** When the queue is full, the client gets `busy` and falls
back to the mailbox. A full queue means the render thread is stuck, and the
systemd watchdog deals with that.
- **Startup.** The server binds under a temporary name, sets the mode and the
group, then renames the socket into place, so it never appears with the
umask's permissions. It removes a stale socket (a file that nothing is
listening on). It never removes a live socket or a file that is not a
socket. `close()` removes the socket only if it is still the one this
process created.
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
as before. The web interface then uses the mailbox.
## Security model
The display runs as root and the web interface as the installing user (see
[PERMISSIONS.md](PERMISSIONS.md)). The socket admits exactly those two, plus
anything else in the group they share:
1. **The directory.** `/run/ledmatrix` is created by `RuntimeDirectory=ledmatrix`
in `ledmatrix.service` (#687): root-owned, `0755`, on tmpfs, and removed
when the display stops. Under an older unit, the display creates the
directory itself as root, as it does for the heartbeat. No installer
change is needed.
2. **The socket file.** The file is `root:<shared group>` with mode `0660`,
and the kernel refuses `connect()` to anyone without write permission on
it. The shared group is the cache directory's group whenever that
directory is group-writable. That is `ledmatrix` on an installed device
(`/var/cache/ledmatrix` is `root:ledmatrix 2775`), and it is the same rule
DiskCache uses for every file the two services share. Otherwise the group
is the project directory's (`get_shared_group_gid()`, which config files
use). With neither, the mode is `0600` and only root can connect.
3. **Peer credentials.** Where the kernel reports them (`SO_PEERCRED`, on
Linux), the server checks every connection again. It accepts root, the
display's own user, or a member of the shared group: the peer's primary
gid, or a supplementary group read from `/proc/<pid>/status`. If `/proc`
is unreadable, it uses the group database. Any other peer gets `forbidden`
and is disconnected. This covers a socket mode that someone loosened by
hand.
The commands are deliberately narrow. Stage 1 can start or stop on-demand
display and read its state, which anyone who can reach the web UI can already
do. Nothing on the socket runs a shell, writes a file, or names a path.
**Development.** A display that is not root and cannot write to
`/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the
socket at `$TMPDIR/ledmatrix-<uid>/control.sock`. That directory is private
(`0700`), and the server refuses it if another user owns it. The web
interface, run by the same user, looks there after `/run/ledmatrix`. The test
suite sets `LEDMATRIX_CONTROL_SOCKET=off` (`test/conftest.py`), so a run on a
device never touches the live display.
## Stage plan
1. **On-demand, with acks (this stage).** Contract, server, client.
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes try the
socket first and report `transport: "socket" | "mailbox"` (plus
`socket_error` on fallback). The mailbox is unchanged, and the plugins that
write it directly (birdnet-go, mqtt-notifications, on-air, pomodoro-timer)
keep working.
2. **Commands that are restarts or polls today.**
- `brightness.set`, transient and with no `config.json` write.
- `plugin.reload`, which replaces the `restart_required` answer from #688
with a live reload of the updated plugin on the render thread.
- `config.reload`, which applies a saved config without waiting for the 2 s
mtime poll and acks which sections changed.
- The dwell sleep and the static screen's 1 s frame sleep wait on the
queue instead of sleeping, so a command lands within milliseconds on
every kind of screen. Under WSL, with a static plugin on screen, a stop
takes 1.0 s by either path today.
3. **A state stream.** A `subscribe` command that keeps the connection open
and pushes events: mode changes, on-demand state, plugin runtime state and
the heartbeat. It replaces the polled `display_current_state`,
`plugin_runtime_snapshot` (#690) and `display-heartbeat.json` (#687) for
readers that hold a connection. The web interface relays it to its
existing SSE stream. The files remain for one release for older readers.
4. **Retire the mailboxes.** After a release in which every device has had the
socket, the web interface stops writing `display_on_demand_request`, and
the display stops polling it, logging the plugins that still write it so
they can move to an in-process `request_display()`. The other cache keys
used as messages (`plugin_error_clear_request` and the remaining
`display_*` keys) move to the socket or to tmpfs.
## Checking it on a device
```bash
ls -l /run/ledmatrix/control.sock # srw-rw---- root ledmatrix
sudo journalctl -u ledmatrix | grep "Control socket"
curl -s -X POST localhost:5000/api/v3/display/on-demand/start \
-H 'Content-Type: application/json' -d '{"plugin_id":"clock","duration":20}'
# ... "transport": "socket"
```
If the response says `"transport": "mailbox"`, `socket_error` gives the
reason. `no_socket` means the display is stopped or predates the socket.
`refused` usually means the web user is not in the socket's group, which
takes effect when the web service restarts after the user is added.
+2
View File
@@ -29,6 +29,8 @@ in again (services pick them up on restart).
| `assets/` | web user | dirs `755`, files `644` | Root writes downloaded logos regardless |
| `/var/cache/ledmatrix/` | `root:ledmatrix` | `2775` (setgid) | Shared cache: see below |
| Cache files | creator : `ledmatrix` | `660` | |
| `/run/ledmatrix/` | `root` | `755` | tmpfs; `RuntimeDirectory=` in `ledmatrix.service`, removed when the display stops |
| `/run/ledmatrix/control.sock` | `root` : cache directory's group (`ledmatrix`) | `660` | The display's control socket; only root and that group can connect. See [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md#security-model) |
| `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them |
| `/etc/sudoers.d/ledmatrix_web`, `ledmatrix_wifi` | `root` | `440` | |
+1
View File
@@ -72,6 +72,7 @@ Going deeper:
## Contributing to LEDMatrix itself
- [ARCHITECTURE.md](ARCHITECTURE.md) — processes, display loop, plugin system, web UI; where to start reading
- [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md) — the display's control socket: protocol, security model, stage plan
- [DEVELOPMENT.md](DEVELOPMENT.md) — environment setup
- [HOW_TO_RUN_TESTS.md](HOW_TO_RUN_TESTS.md) — running the test suite
- [MULTI_ROOT_WORKSPACE_SETUP.md](MULTI_ROOT_WORKSPACE_SETUP.md) — multi-repo workspace
+15 -2
View File
@@ -447,13 +447,23 @@ Request a specific plugin to display on-demand.
"mode": "nfl_live",
"duration": 45,
"pinned": true,
"service": { "active": true, "returncode": 0, "stdout": "", "stderr": "" }
"service": { "active": true, "returncode": 0, "stdout": "", "stderr": "" },
"transport": "socket"
}
}
```
`service` is `null` when `start_service` is false.
`transport` says how the request reached the display: `"socket"` means the
display's control socket acknowledged it (it is queued for the render thread;
see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), `"mailbox"` means it was
written to the cache mailbox the display polls, as before the socket existed.
With `"mailbox"`, `socket_error` gives the reason the socket was not used
(`no_socket` when the display is stopped or predates the socket, `timeout`,
`refused`, `busy`, ...). Either way the request is applied the same way;
`request_id` is the same id in both.
### Stop On-Demand Display
**POST** `/api/v3/display/on-demand/stop`
@@ -476,11 +486,14 @@ Stop the current on-demand display.
"status": "success",
"data": {
"request_id": "uuid-here",
"service": null
"service": null,
"transport": "socket"
}
}
```
`transport` and `socket_error` are as for start.
---
## Plugins