Files
LEDMatrix/scripts/utils/apply_dns_single_request.sh
T
ChuckandClaude Opus 5 59997594ac test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths

Two shared-state problems, both of which show up as a permanently dirty
checkout or an unreproducible test failure.

test_display_dirty_tracking.py builds a real DisplayManager, whose
_snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web
UI reads. Every pytest process on the machine shares that one file, so two
concurrent runs -- CI shards, a second worktree, an agent running the suite
alongside -- overwrite each other's snapshot and the mtime assertions stop
meaning anything. The module fixture now points it at a session-unique temp
path; the individual tests that care still override it further.

To be clear about what this does and does not fix: this is a real shared-path
hazard, but it is NOT the cause of the intermittent 15-test failure in that
module. That turned out to be the emulator's fixed TCP port, fixed in the
follow-up commit. This change stands on its own merits.

web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
and data/operation_history.json as the web interface runs, into a directory
that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig
that ever opened the web UI -- and every test run that constructs the app --
left three untracked files behind and a permanently dirty `git status`. Only
data/.gitkeep is tracked under data/, so the negation keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop the emulator binding a fixed port, so concurrent runs can't collide

This is the cause of the intermittent full-suite failures we have been chasing:
runs of identical code landing anywhere between 100 and 130 failures, while
every implicated test passed in isolation.

Six test modules set EMULATOR=true and build a real DisplayManager. The repo's
emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to
serve the dev preview. That port is a machine-wide singleton, so a second pytest
process -- a CI shard, another worktree, an agent running the suite alongside --
loses the bind. RGBMatrix construction then raises, DisplayManager catches it and
falls back to `self.matrix = None`, and every test that subsequently touches the
matrix dies with

    AttributeError: 'NoneType' object has no attribute 'SwapOnVSync'

which names neither a port nor a socket, and points at the wrong file entirely.
Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its
matrix-touching tests fail together or not at all -- the 15-test swing that made
the totals look random.

Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process
and running test_display_dirty_tracking.py:

    without this change    15 failed, 6 passed
    with this change       21 passed

The "raw" adapter renders in memory and binds nothing. Only display_adapter is
overridden, in a throwaway config written per pytest process; the repo's
emulator_config.json is untouched and `run.py -e` still opens the browser
preview on 8888. Nothing in the suite referenced the adapter, and the tests
wrap SwapOnVSync on the matrix object itself, so they are indifferent to what
sits underneath. allow_adapter_fallback is forced off -- falling back would
land us on the browser adapter and its fixed port, which is the whole problem.

CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set
to an absolute path: the previous behaviour depended on where pytest was invoked
from, and silently wrote a default config into whatever directory that was.

Verified no regressions: full suite on this branch and with origin/main's
versions of the touched files, same machine, back to back -- 115 failed /
4347 passed on both sides, zero failures unique to either. That 115 is the
pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell
scripts, Linux-only binaries); CI on Linux remains authoritative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: mark the shell entry points executable

Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh`
fails with "Permission denied" and only works if you know to prefix `bash`.
That one matters most: the web UI's own error hint, added in #560, tells users
to run exactly that path when a system action fails for want of passwordless
sudo, and following that instruction verbatim did not work.

All eleven carry a shebang and are invoked directly, never sourced. The two
sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately
left non-executable, which is what distinguishes a library from an entry point.

Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied
with `git update-index --chmod=+x` because this checkout is on Windows, where
core.fileMode is off and the working-tree bit is not tracked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:45:13 -04:00

95 lines
3.9 KiB
Bash
Executable File

#!/bin/bash
#
# Add `options single-request` to the system resolver configuration.
#
# glibc's getaddrinfo() sends the A and AAAA queries for a name in
# parallel on one socket. Some routers answer the A query and drop the
# AAAA one, so the resolver waits out its full timeout -- about five
# seconds -- before returning an address that was already available.
# Disabling IPv6 in the kernel does not help: the resolver still asks.
#
# `single-request` makes it send the two queries one after the other,
# which those routers answer correctly. Anything on the matrix that
# calls an external API pays that five seconds per lookup otherwise, and
# a Starlark app with a render timeout will simply fail instead.
#
# Idempotent, and safe to run on a machine that does not need it. Run by
# ledmatrix-dns-fix.service on every boot, because whatever manages
# resolv.conf regenerates it and drops the option again.
#
# Usage: sudo ./scripts/utils/apply_dns_single_request.sh
set -eu
OPTION="options single-request"
RESOLVCONF_TAIL="/etc/resolvconf/resolv.conf.d/tail"
RESOLV_CONF="/etc/resolv.conf"
log() { echo "[dns-single-request] $*"; }
already_applied() {
grep -qs "^${OPTION}\$" "$1"
}
# resolvconf regenerates /etc/resolv.conf from these fragments, so the
# tail file is the only place an addition survives. Prefer it when the
# directory exists, whether or not resolvconf has run yet.
if [ -d "$(dirname "$RESOLVCONF_TAIL")" ]; then
if already_applied "$RESOLVCONF_TAIL"; then
log "already present in $RESOLVCONF_TAIL"
else
echo "$OPTION" >> "$RESOLVCONF_TAIL"
log "added to $RESOLVCONF_TAIL"
fi
# Only a missing resolvconf is ignorable. If it is present and the
# regeneration fails, /etc/resolv.conf still lacks the option, and
# reporting success would be a lie.
if command -v resolvconf >/dev/null 2>&1; then
if ! resolvconf -u; then
log "resolvconf -u failed; $RESOLV_CONF was not regenerated"
exit 1
fi
fi
fi
# systemd-resolved owns its stub file and rewrites anything appended to it,
# and `single-request` is a glibc resolv.conf option with no resolved.conf
# equivalent -- so there is nothing this script can do here. Exit non-zero:
# the unit would otherwise record success while the workaround is inactive,
# which is the failure mode this whole script exists to avoid.
if [ -L "$RESOLV_CONF" ] && readlink -f "$RESOLV_CONF" | grep -q "systemd"; then
log "$RESOLV_CONF is managed by systemd-resolved."
log "'options single-request' is a glibc resolv.conf option and has no"
log "resolved.conf equivalent, so it cannot be applied on this host."
log "If external API calls are slow, the workaround is to stop using the"
log "systemd-resolved stub (see 'man systemd-resolved', NSS/resolv.conf modes)."
exit 1
fi
if already_applied "$RESOLV_CONF"; then
log "already present in $RESOLV_CONF"
exit 0
fi
if [ ! -w "$RESOLV_CONF" ] && [ -e "$RESOLV_CONF" ]; then
log "cannot write $RESOLV_CONF (run with sudo?)"
exit 1
fi
# A NetworkManager-generated resolv.conf is regenerated on every connection
# change, not only at boot -- and this unit is oneshot with RemainAfterExit,
# so it will not re-run within the same boot to put the option back. Say so
# rather than implying the fix is permanent. Nothing is silently swallowed:
# the append below still happens and still works until the next renewal.
if grep -qs "Generated by NetworkManager" "$RESOLV_CONF" \
&& [ ! -d "$(dirname "$RESOLVCONF_TAIL")" ]; then
log "NOTE: $RESOLV_CONF is generated by NetworkManager and has no"
log "resolvconf tail directory to write to. The option is being added, but"
log "NetworkManager will drop it on the next connection renewal, and this"
log "unit does not run again until the next boot. If lookups go slow again"
log "before a reboot, re-run this script."
fi
echo "$OPTION" >> "$RESOLV_CONF"
log "added to $RESOLV_CONF"