mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 22:35:08 +00:00
c4c46d3ba76e2d9443551324e1af65db240c9b98
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
59997594ac |
test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths Two shared-state problems, both of which show up as a permanently dirty checkout or an unreproducible test failure. test_display_dirty_tracking.py builds a real DisplayManager, whose _snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web UI reads. Every pytest process on the machine shares that one file, so two concurrent runs -- CI shards, a second worktree, an agent running the suite alongside -- overwrite each other's snapshot and the mtime assertions stop meaning anything. The module fixture now points it at a session-unique temp path; the individual tests that care still override it further. To be clear about what this does and does not fix: this is a real shared-path hazard, but it is NOT the cause of the intermittent 15-test failure in that module. That turned out to be the emulator's fixed TCP port, fixed in the follow-up commit. This change stands on its own merits. web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json and data/operation_history.json as the web interface runs, into a directory that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig that ever opened the web UI -- and every test run that constructs the app -- left three untracked files behind and a permanently dirty `git status`. Only data/.gitkeep is tracked under data/, so the negation keeps it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: stop the emulator binding a fixed port, so concurrent runs can't collide This is the cause of the intermittent full-suite failures we have been chasing: runs of identical code landing anywhere between 100 and 130 failures, while every implicated test passed in isolation. Six test modules set EMULATOR=true and build a real DisplayManager. The repo's emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to serve the dev preview. That port is a machine-wide singleton, so a second pytest process -- a CI shard, another worktree, an agent running the suite alongside -- loses the bind. RGBMatrix construction then raises, DisplayManager catches it and falls back to `self.matrix = None`, and every test that subsequently touches the matrix dies with AttributeError: 'NoneType' object has no attribute 'SwapOnVSync' which names neither a port nor a socket, and points at the wrong file entirely. Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its matrix-touching tests fail together or not at all -- the 15-test swing that made the totals look random. Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process and running test_display_dirty_tracking.py: without this change 15 failed, 6 passed with this change 21 passed The "raw" adapter renders in memory and binds nothing. Only display_adapter is overridden, in a throwaway config written per pytest process; the repo's emulator_config.json is untouched and `run.py -e` still opens the browser preview on 8888. Nothing in the suite referenced the adapter, and the tests wrap SwapOnVSync on the matrix object itself, so they are indifferent to what sits underneath. allow_adapter_fallback is forced off -- falling back would land us on the browser adapter and its fixed port, which is the whole problem. CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set to an absolute path: the previous behaviour depended on where pytest was invoked from, and silently wrote a default config into whatever directory that was. Verified no regressions: full suite on this branch and with origin/main's versions of the touched files, same machine, back to back -- 115 failed / 4347 passed on both sides, zero failures unique to either. That 115 is the pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell scripts, Linux-only binaries); CI on Linux remains authoritative. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: mark the shell entry points executable Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh` fails with "Permission denied" and only works if you know to prefix `bash`. That one matters most: the web UI's own error hint, added in #560, tells users to run exactly that path when a system action fails for want of passwordless sudo, and following that instruction verbatim did not work. All eleven carry a shebang and are invoked directly, never sourced. The two sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately left non-executable, which is what distinguishes a library from an entry point. Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied with `git update-index --chmod=+x` because this checkout is on Windows, where core.fileMode is off and the working-tree bit is not tracked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e23f1f45d3 |
feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of patches and services built while running this project on Starlark apps under MQTT control. Its patches are whole-file copies taken against an older tree, so applying them as written would revert #523's frame pacing, #534's display() bool returns and the GitHub token masking in plugins_manager.js. Three of its claimed fixes are already in main, and its api_v3 Starlark routes are #535's. What follows is the rest -- verified against current code, and reimplemented where the patch's approach did not hold up. **On-demand display.** `pinned` reached the controller from the API, was stored on it and republished in the status payload, but never narrowed the rotation -- a pinned request still cycled every mode its plugin owns. Right for a sports plugin, whose modes are views of one subject; wrong for a plugin whose modes are unrelated, which is every Starlark app. Now honoured, and it survives a restart. Restarting while on-demand was active loaded *only* the on-demand plugin, so normal rotation had nothing to return to for the life of the process -- and a restart mid-session is routine, since that is how an update is applied. The panel came back cycling one plugin's modes with no way out but clearing the cache by hand. Every enabled plugin loads now; on-demand still resumes on its saved mode. Stop requests are exempt from the duplicate guards on purpose, so that a second click stops a mode a race left running -- which means consuming the mailbox is the only thing that ends one. It was never consumed, so the same stop was re-read and re-processed on every poll, forever. Both paths now share one compare-before-delete helper. **Starlark rendering.** `extract_schema` parsed the source with a regex, which can only see option lists written out literally: an app whose dropdown is filled from a live API call inside `get_schema()` came back empty, and the config form offered nothing to pick. Now runs `pixlet schema`, which executes the app, and falls back to the parser when Pixlet is absent, too old for the subcommand, or the app fails to run. The third-party patch replaced the parser outright and hardcoded /usr/local/bin/pixlet; this keeps the fallback and the binary search. A `|` in a config value was dropped by a shell-metacharacter filter, though the command is a list with no shell involved -- and apps do use it as a separator inside one value. The key went missing silently and the app rendered its own "not configured" screen with nothing to say why. And a 0-byte render was reported as success: Pixlet exits 0 and writes nothing when an app has no content, which read downstream as a working app drawing a black panel. **Starlark display.** `display()` ignored the mode it was called with, so a specific app could not be addressed. It now accepts `display_mode` -- which is the whole mechanism, since the controller inspects the signature before passing it. Found while there: `_select_next_app` ran only while `current_app` was unset, so with several apps installed the first was picked once and shown forever while the rest were rendered on schedule and never displayed. And `enable_scrolling` was missing, so multi-frame apps were called once per rotation slot and never advanced past frame one. **GET /api/v3/display/modes.** Every mode that can be requested on-demand, with the plugin that owns it. Nothing exposed this, so anything driving the display from outside the web UI read each plugin's manifest.json off disk and reimplemented PluginManager's fallbacks. It also triggers discovery, which is otherwise lazy and normally happens because a person opened the dashboard. **integrations/mqtt_bridge.** Home Assistant control over MQTT Discovery: a mode select, a stop button, power, brightness. Rewritten against the API rather than the filesystem, so it needs no read access to config.json and cannot drift from the web UI. paho-mqtt 2.x VERSION2, TLS, an availability topic that is also the last will, and secrets from the environment. **Two opt-in extras.** A DNS single-request unit, for glibc's parallel A/AAAA lookup stalling ~5s per name on routers that answer only the A query -- which makes any plugin calling an external API slow and Starlark apps, which have a render timeout, fail outright. And a Pixlet config editor: a script you run and Ctrl+C rather than the third-party version's always-on unauthenticated Flask service, since it stops the display for the length of a session. Neither is installed by default. Long Starlark app names now wrap instead of overflowing their card. 115 new tests across 5 files. Also unblocked test_starlark_display_contract.py, which was silently skipping wherever fcntl is absent. Whole suite: no new failures against main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(mqtt_bridge): the five issues Codacy flagged on this branch All in the new bridge, all real: * requests floor was 2.31.0, which carries CVE-2024-35195, CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which is what the project's own requirements.txt already pins. * `import time` was never used. * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential. It is the "no password configured" default; marked nosec B105, the convention used elsewhere in the repo. Also dropped an unused `build_app` from the display-modes test imports. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: the review findings on this PR Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and is answered below. **One bad config section blanked the whole mode list.** `/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`, so a non-dict under a plugin id -- a shape DisplayController already guards, so it happens -- raised AttributeError mid-loop and answered 500 with no modes at all. Every MQTT bridge entity is built from that list. Now skipped with a warning. **The DNS scripts reported success they had not earned.** Three separate paths: `resolvconf -u` failing was swallowed by `|| true`; the systemd-resolved branch exited 0 without applying anything, so the oneshot unit recorded success while the workaround was inactive; and the installer's `|| echo` turned a failed start into "installation complete." with exit 0. All three now fail loudly. `single-request` is a glibc resolv.conf option with no resolved.conf equivalent, so on those hosts the honest answer is that it cannot be applied. A NetworkManager-generated resolv.conf is regenerated on connection changes, not only at boot, and the unit is oneshot with RemainAfterExit -- so the option can vanish mid-boot with nothing to put it back. Now detected and stated plainly rather than implied to be permanent. **`Before=` does not order a manual restart.** It only orders units already in the same transaction, so `systemctl restart ledmatrix` could bypass the fix. install_dns_fix.sh now writes a ledmatrix.service drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround failing should not stop the display. **The Pixlet editor's `--lan` is gone.** `pixlet serve` has no authentication, and a printed warning is not access control. Loopback only, with the SSH port-forward in the header where the flag used to be documented -- SSH does the authenticating and nothing is left listening. **The MQTT example config now defaults to TLS** on 8883. The installer copies it verbatim, and without TLS the broker password and every command cross the network in cleartext. A plaintext broker is still supported and documented, and the bridge warns once at startup when a password is configured without TLS. **Not taken: "the upstream Pixlet CLI has no `schema` subcommand."** Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh` installs `tronbyt/pixlet`, whose `cmd/schema.go` is `schema [PATH]` -> JSON on stdout, built on `runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is exactly what extract_schema_via_pixlet calls. A binary without the subcommand exits non-zero and falls back to the source parser, which is already covered by a test. **CodeQL stack-trace exposure: not taken either.** I removed `details` first and that broke test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception, which enforces `describe_exception` across all ~75 handlers -- written because a device with failing storage answered "see logs for details" from the log viewer itself. describe_exception redacts credentials; the trade-off is the project's and is already made. Restored, with the reasoning in a comment. 11 new tests. Whole suite: no new failures against main, 4127 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |