mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 22:35:08 +00:00
feat(ipc): display control socket, stage 1: on-demand with acks
The display process now serves a Unix socket, /run/ledmatrix/control.sock, carrying versioned newline-delimited JSON commands that are acknowledged. Stage 1 moves on-demand start/stop (plus status, hello and ping) onto it; the cache-file mailbox stays as the fallback for one release. - src/ipc/contract.py: typed request/response envelopes, command args, error codes, NDJSON framing with a 64 KiB limit, socket path rules. - src/ipc/server.py: threaded server owned by the display. Handlers only queue onto a bounded queue and ack with the request id; the render thread drains it where it reads the mailbox. Bounded clients, timeouts, garbage/oversize/disconnect handling; 0660 socket in the cache dir's group plus SO_PEERCRED checks; skips cleanly on Windows or when off. - src/ipc/client.py: one short-timeout request; any failure raises ControlError(reason). - api_v3/display.py: on-demand start/stop try the socket, fall back to the mailbox exactly as before, and report transport/socket_error. - display_controller.py: start/close the server; the mailbox handler body is extracted into _handle_on_demand_request and shared by both paths. - docs/IPC_CONTROL_SOCKET.md: protocol, security model, stage plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
+10
-4
@@ -41,7 +41,8 @@ each other. They share three things:
|
||||
|
||||
| State | Where | Written by | Read by |
|
||||
|---|---|---|---|
|
||||
| On-demand request | cache `display_on_demand_request` | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py) | display: `_poll_on_demand_requests()` |
|
||||
| On-demand command | control socket `/run/ledmatrix/control.sock` ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py), via [`src/ipc/client.py`](../src/ipc/client.py) | display: [`src/ipc/server.py`](../src/ipc/server.py) acks; the render thread applies it in `_poll_on_demand_requests()` |
|
||||
| On-demand request (fallback) | cache `display_on_demand_request` | web, when the socket fails; four plugins write it directly | display: `_poll_on_demand_requests()` |
|
||||
| On-demand state | cache `display_on_demand_state` | display: `_publish_on_demand_state()` | web: `/api/v3/display/on-demand/status` |
|
||||
| Current screen | cache `display_current_state` | display | web: `/api/v3/display/current-status` |
|
||||
| Plugin errors | cache `plugin_error_snapshot` | display: `ErrorSnapshotPublisher` ([`src/error_aggregator.py`](../src/error_aggregator.py)) | web: `read_error_report()` for `/api/v3/errors/*` |
|
||||
@@ -55,9 +56,14 @@ each other. They share three things:
|
||||
| Render-loop heartbeat | `/run/ledmatrix/display-heartbeat.json` (tmpfs) | display: the render thread, via [`display_watchdog`](../src/display_watchdog.py) | web: `/api/v3/health` (`checks.display_loop`); the update health check |
|
||||
|
||||
The on-demand start route starts `ledmatrix.service` when it is not running
|
||||
(`start_service`, on by default) but never restarts a running one: the display
|
||||
reads the mailbox every `ON_DEMAND_POLL_INTERVAL` (0.25s), from its dwell
|
||||
sleep, its render loops and Vegas's interrupt check as well as the main loop.
|
||||
(`start_service`, on by default) but never restarts a running one. The routes
|
||||
send the command over the display's control socket and get an ack; when that
|
||||
fails (a stopped display, one older than the socket) they write the mailbox
|
||||
instead, which the display reads every `ON_DEMAND_POLL_INTERVAL` (0.25s), from
|
||||
its dwell sleep, its render loops and Vegas's interrupt check as well as the
|
||||
main loop. Both ways end in the same handler, `_handle_on_demand_request()`.
|
||||
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
|
||||
for the protocol, the permission model and the plan to retire the mailboxes.
|
||||
|
||||
### Web and display processes: who runs plugins
|
||||
|
||||
|
||||
@@ -0,0 +1,250 @@
|
||||
# Control socket (web → display)
|
||||
|
||||
The display process serves a Unix socket that the web interface uses to send
|
||||
it commands and get an answer back. It replaces the cache-file "mailboxes" on
|
||||
the SD card one command at a time. Stage 1, described here, carries on-demand
|
||||
start/stop/status. The file mailbox stays as a fallback for one release.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Socket | `/run/ledmatrix/control.sock` (tmpfs) |
|
||||
| Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` |
|
||||
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop` |
|
||||
| Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it |
|
||||
| Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it |
|
||||
|
||||
## Why
|
||||
|
||||
Before the socket, the web interface sent commands by writing a cache key
|
||||
(`display_on_demand_request`) that the display read every 0.25 s.
|
||||
|
||||
- **No acknowledgement.** The route answered "success" once the file was
|
||||
written, whether or not a display was running to read it.
|
||||
- **Lost requests.** The display had to read the request and then delete it.
|
||||
A request written between those two steps could be thrown away (see
|
||||
`_consume_on_demand_request`). The cache has no atomic claim to prevent it.
|
||||
- **Fragile.** Each channel repeated its own permission, atomic-write,
|
||||
staleness and in-memory-cache rules. Two of them caused bugs: a `memory_ttl`
|
||||
bug ignored every on-demand request after the first for an hour, and a
|
||||
stopped display was still reported as "active" for two minutes.
|
||||
|
||||
The socket answers every command, carries one request per message (so nothing
|
||||
can overwrite it), and belongs to the display process. If the display is not
|
||||
running, the socket does not exist, and the web interface knows right away.
|
||||
|
||||
## Protocol (version 1)
|
||||
|
||||
**Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at
|
||||
most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with
|
||||
`ensure_ascii`, so a newline never appears inside a message. A connection
|
||||
can carry several requests. Each request gets exactly one response, in order.
|
||||
|
||||
**Request**
|
||||
|
||||
```json
|
||||
{"v": 1, "id": "5f0c…", "cmd": "on_demand.start",
|
||||
"args": {"plugin_id": "clock", "mode": null, "duration": 30, "pinned": false}}
|
||||
```
|
||||
|
||||
- `v` is the protocol version.
|
||||
- `id` is a printable string of 1-128 characters. It is echoed back in the
|
||||
response, and for on-demand commands it is also the on-demand `request_id`.
|
||||
- `cmd` is a command name.
|
||||
- `args` is an object. It may be omitted when a command takes no arguments.
|
||||
|
||||
**Response**
|
||||
|
||||
```json
|
||||
{"v": 1, "id": "5f0c…", "ok": true, "result": {"accepted": true, "request_id": "5f0c…", "queued": 1}}
|
||||
{"v": 1, "id": "5f0c…", "ok": false, "error": {"code": "busy", "message": "…"}}
|
||||
```
|
||||
|
||||
`id` is `null` only when the request could not be parsed far enough to have
|
||||
one. Clients branch on `error.code`, never on the message text.
|
||||
|
||||
**Commands**
|
||||
|
||||
| `cmd` | `args` | `result` | Kind |
|
||||
|---|---|---|---|
|
||||
| `hello` | `{versions: [int], client?: str}` | `{version, versions, commands, max_message_bytes, server}` | answered directly |
|
||||
| `ping` | — | `{pong: true}` | answered directly |
|
||||
| `on_demand.start` | `{plugin_id?, mode?, duration?, pinned?}` (at least one of `plugin_id` and `mode`) | ack | queued |
|
||||
| `on_demand.stop` | — | ack | queued |
|
||||
| `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly |
|
||||
|
||||
`duration` is a number of seconds, or a numeric string. `0`, `null` or `""`
|
||||
mean "until stopped". `pinned` must be a real boolean: the REST route has
|
||||
already converted strings like `"false"` before it sends the command. The
|
||||
`on_demand` object in `on_demand.status` is the same dict the display
|
||||
publishes to `display_on_demand_state`.
|
||||
|
||||
**Acknowledgements.** A queued command is *accepted*, not *done*.
|
||||
`{"accepted": true, "request_id": …}` means the command is waiting in the
|
||||
render thread's queue, and the render thread will apply it at its next
|
||||
on-demand check. That is within one frame on a scrolling screen, 0.25 s
|
||||
during a dwell, and up to 1 s on a static screen, whose frame loop sleeps a
|
||||
second between frames. Except on a scrolling screen, where the mailbox waits
|
||||
up to 0.25 s, these are the mailbox's delays too: stage 1 adds
|
||||
acknowledgements, not speed. Any outcome is published as before
|
||||
(`display_on_demand_state`, and `status`/`error` for a bad plugin or mode),
|
||||
and it can be read with `on_demand.status`.
|
||||
|
||||
**Versions.** Every request carries `v`. For any command except `hello`, a
|
||||
`v` the display does not speak gets `unsupported_version`. `hello` is checked
|
||||
by its `versions` list instead, and its result names the highest version both
|
||||
sides share, so a client can find out what a display supports before it
|
||||
relies on anything newer. Stage 1's client sends `v: 1` and falls back to the
|
||||
mailbox when the display refuses it. It does not send `hello` first, which
|
||||
saves a round trip.
|
||||
|
||||
**Error codes:** `bad_json`, `bad_request`, `message_too_large`,
|
||||
`unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full,
|
||||
or too many connections), `forbidden` (peer credentials refused), `internal`.
|
||||
|
||||
Try it on a device:
|
||||
|
||||
```bash
|
||||
python3 - <<'EOF'
|
||||
from src.ipc import client # run from the project directory
|
||||
print(client.on_demand_status())
|
||||
EOF
|
||||
```
|
||||
|
||||
## How the display applies a command
|
||||
|
||||
The server's threads never touch rendering. A connection thread parses the
|
||||
request, validates it against the contract, and then does one of two things:
|
||||
|
||||
- For a command that changes the panel, it puts a `QueuedCommand` on a
|
||||
bounded queue (16 entries) and answers with the ack.
|
||||
- For a query, it answers from a status snapshot the display provides
|
||||
(`DisplayController._control_status`). The snapshot only reads attributes.
|
||||
|
||||
The render thread drains the queue in `_poll_on_demand_requests()`, the same
|
||||
place it reads the mailbox, and hands each command to
|
||||
`_handle_on_demand_request()`, which is the mailbox's own handler. The two
|
||||
paths share all of their code: activation, the processed-id guard, error
|
||||
publishing, and resuming the rotation afterwards. The 0.25 s floor on the
|
||||
mailbox read does not apply to the queue, because draining it costs no disk
|
||||
read. A queued command also lets `_service_pending_changes()` skip its own
|
||||
floor, so a long scrolling screen or a Vegas iteration takes the command at
|
||||
its next frame.
|
||||
|
||||
**Exactly once.** A command and a mailbox write for the same request share
|
||||
one `request_id`. If the client times out after the display queued the
|
||||
command and then also writes the mailbox, the display processes the request
|
||||
once. The existing `on_demand_request_id` and processed-id checks drop the
|
||||
second copy.
|
||||
|
||||
## Robustness
|
||||
|
||||
All of this runs inside the display process, so nothing a client does may
|
||||
block the render loop or crash it:
|
||||
|
||||
- **Bounded connections.** Each connection gets its own daemon thread, with
|
||||
at most 8 at once. One more is answered `busy` and closed.
|
||||
- **Timeouts.** Each read and write times out after 2 s. A message must
|
||||
arrive whole within 5 s of its first byte. An idle connection is closed
|
||||
after 10 s. A slow or stuck client costs one thread for a few seconds.
|
||||
- **Malformed input.** A line that is not JSON gets `bad_json`, and the
|
||||
connection carries on. A line longer than 64 KiB gets `message_too_large`,
|
||||
and the connection is closed, because the next message boundary cannot be
|
||||
found. A client that disconnects mid-message is dropped silently. No
|
||||
exception from a handler leaves the connection thread.
|
||||
- **Full queue.** When the queue is full, the client gets `busy` and falls
|
||||
back to the mailbox. A full queue means the render thread is stuck, and the
|
||||
systemd watchdog deals with that.
|
||||
- **Startup.** The server binds under a temporary name, sets the mode and the
|
||||
group, then renames the socket into place, so it never appears with the
|
||||
umask's permissions. It removes a stale socket (a file that nothing is
|
||||
listening on). It never removes a live socket or a file that is not a
|
||||
socket. `close()` removes the socket only if it is still the one this
|
||||
process created.
|
||||
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
|
||||
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
|
||||
as before. The web interface then uses the mailbox.
|
||||
|
||||
## Security model
|
||||
|
||||
The display runs as root and the web interface as the installing user (see
|
||||
[PERMISSIONS.md](PERMISSIONS.md)). The socket admits exactly those two, plus
|
||||
anything else in the group they share:
|
||||
|
||||
1. **The directory.** `/run/ledmatrix` is created by `RuntimeDirectory=ledmatrix`
|
||||
in `ledmatrix.service` (#687): root-owned, `0755`, on tmpfs, and removed
|
||||
when the display stops. Under an older unit, the display creates the
|
||||
directory itself as root, as it does for the heartbeat. No installer
|
||||
change is needed.
|
||||
2. **The socket file.** The file is `root:<shared group>` with mode `0660`,
|
||||
and the kernel refuses `connect()` to anyone without write permission on
|
||||
it. The shared group is the cache directory's group whenever that
|
||||
directory is group-writable. That is `ledmatrix` on an installed device
|
||||
(`/var/cache/ledmatrix` is `root:ledmatrix 2775`), and it is the same rule
|
||||
DiskCache uses for every file the two services share. Otherwise the group
|
||||
is the project directory's (`get_shared_group_gid()`, which config files
|
||||
use). With neither, the mode is `0600` and only root can connect.
|
||||
3. **Peer credentials.** Where the kernel reports them (`SO_PEERCRED`, on
|
||||
Linux), the server checks every connection again. It accepts root, the
|
||||
display's own user, or a member of the shared group: the peer's primary
|
||||
gid, or a supplementary group read from `/proc/<pid>/status`. If `/proc`
|
||||
is unreadable, it uses the group database. Any other peer gets `forbidden`
|
||||
and is disconnected. This covers a socket mode that someone loosened by
|
||||
hand.
|
||||
|
||||
The commands are deliberately narrow. Stage 1 can start or stop on-demand
|
||||
display and read its state, which anyone who can reach the web UI can already
|
||||
do. Nothing on the socket runs a shell, writes a file, or names a path.
|
||||
|
||||
**Development.** A display that is not root and cannot write to
|
||||
`/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the
|
||||
socket at `$TMPDIR/ledmatrix-<uid>/control.sock`. That directory is private
|
||||
(`0700`), and the server refuses it if another user owns it. The web
|
||||
interface, run by the same user, looks there after `/run/ledmatrix`. The test
|
||||
suite sets `LEDMATRIX_CONTROL_SOCKET=off` (`test/conftest.py`), so a run on a
|
||||
device never touches the live display.
|
||||
|
||||
## Stage plan
|
||||
|
||||
1. **On-demand, with acks (this stage).** Contract, server, client.
|
||||
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes try the
|
||||
socket first and report `transport: "socket" | "mailbox"` (plus
|
||||
`socket_error` on fallback). The mailbox is unchanged, and the plugins that
|
||||
write it directly (birdnet-go, mqtt-notifications, on-air, pomodoro-timer)
|
||||
keep working.
|
||||
2. **Commands that are restarts or polls today.**
|
||||
- `brightness.set`, transient and with no `config.json` write.
|
||||
- `plugin.reload`, which replaces the `restart_required` answer from #688
|
||||
with a live reload of the updated plugin on the render thread.
|
||||
- `config.reload`, which applies a saved config without waiting for the 2 s
|
||||
mtime poll and acks which sections changed.
|
||||
- The dwell sleep and the static screen's 1 s frame sleep wait on the
|
||||
queue instead of sleeping, so a command lands within milliseconds on
|
||||
every kind of screen. Under WSL, with a static plugin on screen, a stop
|
||||
takes 1.0 s by either path today.
|
||||
3. **A state stream.** A `subscribe` command that keeps the connection open
|
||||
and pushes events: mode changes, on-demand state, plugin runtime state and
|
||||
the heartbeat. It replaces the polled `display_current_state`,
|
||||
`plugin_runtime_snapshot` (#690) and `display-heartbeat.json` (#687) for
|
||||
readers that hold a connection. The web interface relays it to its
|
||||
existing SSE stream. The files remain for one release for older readers.
|
||||
4. **Retire the mailboxes.** After a release in which every device has had the
|
||||
socket, the web interface stops writing `display_on_demand_request`, and
|
||||
the display stops polling it, logging the plugins that still write it so
|
||||
they can move to an in-process `request_display()`. The other cache keys
|
||||
used as messages (`plugin_error_clear_request` and the remaining
|
||||
`display_*` keys) move to the socket or to tmpfs.
|
||||
|
||||
## Checking it on a device
|
||||
|
||||
```bash
|
||||
ls -l /run/ledmatrix/control.sock # srw-rw---- root ledmatrix
|
||||
sudo journalctl -u ledmatrix | grep "Control socket"
|
||||
curl -s -X POST localhost:5000/api/v3/display/on-demand/start \
|
||||
-H 'Content-Type: application/json' -d '{"plugin_id":"clock","duration":20}'
|
||||
# ... "transport": "socket"
|
||||
```
|
||||
|
||||
If the response says `"transport": "mailbox"`, `socket_error` gives the
|
||||
reason. `no_socket` means the display is stopped or predates the socket.
|
||||
`refused` usually means the web user is not in the socket's group, which
|
||||
takes effect when the web service restarts after the user is added.
|
||||
@@ -29,6 +29,8 @@ in again (services pick them up on restart).
|
||||
| `assets/` | web user | dirs `755`, files `644` | Root writes downloaded logos regardless |
|
||||
| `/var/cache/ledmatrix/` | `root:ledmatrix` | `2775` (setgid) | Shared cache: see below |
|
||||
| Cache files | creator : `ledmatrix` | `660` | |
|
||||
| `/run/ledmatrix/` | `root` | `755` | tmpfs; `RuntimeDirectory=` in `ledmatrix.service`, removed when the display stops |
|
||||
| `/run/ledmatrix/control.sock` | `root` : cache directory's group (`ledmatrix`) | `660` | The display's control socket; only root and that group can connect. See [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md#security-model) |
|
||||
| `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them |
|
||||
| `/etc/sudoers.d/ledmatrix_web`, `ledmatrix_wifi` | `root` | `440` | |
|
||||
|
||||
|
||||
@@ -72,6 +72,7 @@ Going deeper:
|
||||
## Contributing to LEDMatrix itself
|
||||
|
||||
- [ARCHITECTURE.md](ARCHITECTURE.md) — processes, display loop, plugin system, web UI; where to start reading
|
||||
- [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md) — the display's control socket: protocol, security model, stage plan
|
||||
- [DEVELOPMENT.md](DEVELOPMENT.md) — environment setup
|
||||
- [HOW_TO_RUN_TESTS.md](HOW_TO_RUN_TESTS.md) — running the test suite
|
||||
- [MULTI_ROOT_WORKSPACE_SETUP.md](MULTI_ROOT_WORKSPACE_SETUP.md) — multi-repo workspace
|
||||
|
||||
@@ -447,13 +447,23 @@ Request a specific plugin to display on-demand.
|
||||
"mode": "nfl_live",
|
||||
"duration": 45,
|
||||
"pinned": true,
|
||||
"service": { "active": true, "returncode": 0, "stdout": "", "stderr": "" }
|
||||
"service": { "active": true, "returncode": 0, "stdout": "", "stderr": "" },
|
||||
"transport": "socket"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`service` is `null` when `start_service` is false.
|
||||
|
||||
`transport` says how the request reached the display: `"socket"` means the
|
||||
display's control socket acknowledged it (it is queued for the render thread;
|
||||
see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), `"mailbox"` means it was
|
||||
written to the cache mailbox the display polls, as before the socket existed.
|
||||
With `"mailbox"`, `socket_error` gives the reason the socket was not used
|
||||
(`no_socket` when the display is stopped or predates the socket, `timeout`,
|
||||
`refused`, `busy`, ...). Either way the request is applied the same way;
|
||||
`request_id` is the same id in both.
|
||||
|
||||
### Stop On-Demand Display
|
||||
|
||||
**POST** `/api/v3/display/on-demand/stop`
|
||||
@@ -476,11 +486,14 @@ Stop the current on-demand display.
|
||||
"status": "success",
|
||||
"data": {
|
||||
"request_id": "uuid-here",
|
||||
"service": null
|
||||
"service": null,
|
||||
"transport": "socket"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`transport` and `socket_error` are as for start.
|
||||
|
||||
---
|
||||
|
||||
## Plugins
|
||||
|
||||
Reference in New Issue
Block a user