Compare commits

..
Author SHA1 Message Date
ChuckBuilds 26324e7fc5 Merge branch 'push/version-reporting' into push/compatibility-gate 2026-08-02 20:40:18 -04:00
ChuckBuildsandClaude Opus 5 e26ed29385 fix(version): address review — regex strictness, OSError, stale doc claim
From CodeRabbit on #428, all three valid:

- The module docstring claimed the tag check "runs at release time in
  .github/workflows/release-version-check.yml". That workflow is held back to
  a follow-up PR (the pushing token lacks the `workflow` scope), so the claim
  was false as written. Both files now describe the script as a manual
  pre-flight and say the CI wiring is still to come.

- `\d` also matches non-ASCII decimal digits, which int() happily parses, and
  `\s` matches newlines -- so "##\n3.2.0" read as a version heading. Patterns
  now use [0-9] and [ \t], kept in step across the test and the script, with a
  regression test pinning both behaviours.

- A missing or unreadable CHANGELOG.md raised OSError out of read_text() and
  printed a traceback. In a release gate that reads as "the tooling is
  broken"; it now reports the path and a recovery action and exits 1.

Verified: v3.2.0 passes, a mismatched tag exits 1, and a missing CHANGELOG
exits 1 with the new message instead of a traceback. 5 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
2026-08-02 20:40:06 -04:00
ChuckBuildsandClaude Opus 5 c615d5a3bb fix(store): serialize concurrent installs, and make the lock reentrant
Second bug found while validating the previous commit on hardware.

install_plugin's new set-aside/restore had no lock. The web UI runs Flask
threaded, so a double-clicked Install button gives two threads the same
plugin_id; interleaved, one thread's restore deletes the other's freshly
installed copy. _reinstall_with_rollback already guards exactly this with a
per-plugin lock, and install_plugin needs the same one.

Taking that lock naively deadlocks. _reinstall_with_rollback holds it across
its call to install_plugin, and threading.Lock is not reentrant -- so the
request thread hangs forever on the standard monorepo update path
(update_plugin -> _reinstall_with_rollback -> install_plugin), which is to say
on every plugin update. Verified by reverting to a plain Lock: the regression
test times out after 10s instead of passing.

The per-plugin locks are now RLocks, and install_plugin holds one for its
whole set-aside/install/restore sequence.

Verified on devpi (Pi, Python 3.13.5, real registry and network):
- update_plugin on an up-to-date plugin: True in 5.4s
- update_plugin forced through the full reinstall-with-rollback path:
  True in 13.1s, correct version restored, old copy replaced, no backup
  directories left behind
- install -> reinstall-over-existing -> failed-reinstall-restores: all pass
  against real downloads
- 22 plugins load, no tracebacks, web API and UI 200, steady-state journal
  50 lines/min

791 core unit tests pass, including 2 new concurrency tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
2026-08-02 20:33:24 -04:00
ChuckBuildsandClaude Opus 5 fecd9e1385 fix(store): a failed install must not destroy the plugin it replaced
Found while validating the compatibility gate. `_install_plugin_impl` deletes
the existing plugin directory *before* downloading, so any failure after that
point leaves the user with nothing. `_reinstall_with_rollback` protects the
update path exactly this way; a direct `install_plugin` had no equivalent.

The gate made this reachable in a new way: a plugin whose declared floor
exceeds the running core is now refused *after* the old copy is already gone.
Floors are hand-written and can be over-declared, so the refusal could remove
a plugin that had been working fine on that core.

install_plugin is now a thin wrapper that renames any existing install aside,
delegates to _install_plugin_impl, and restores it on failure -- including
when the implementation raises, which is re-raised after the restore. It is a
pass-through when nothing is installed and when called from
_reinstall_with_rollback, which has already moved the old copy aside; a test
pins that so the two mechanisms cannot start nesting.

The aside name embeds '.standalone-backup-' because
plugin_manager._scan_directory_for_plugins keys on exactly that substring to
skip backups. A different name would have made the backup discoverable as a
duplicate plugin; a test pins that too.

789 core unit tests pass, including 7 new ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
2026-08-02 20:33:24 -04:00
ChuckBuildsandClaude Opus 5 749a6a9028 feat(store): refuse to install a plugin that needs a newer core
`ledmatrix_min_version` was decoration. The loader logged an advisory warning
and continued; the store never compared the core version at all, so a routine
"update" delivered a plugin that could not run. That is the gap phase B6 (the
sports-unification sunset) cannot be done over: deleting a plugin's bundled
fallback while nothing enforces the floor hands un-updated users a scoreboard
that raises ModuleNotFoundError at load and is reported only as one line in
the journal.

The gate lives in install_plugin, after the manifest is on disk and before
dependencies are installed. That is the earliest knowable point -- the
registry carries no compatibility field, so the floor is not visible until
the files are down -- and it is also the chokepoint: _reinstall_with_rollback
calls install_plugin, so a refused *update* restores the version the user
already had, for free.

Floor resolution and the comparison move to src/plugin_system/compatibility.py,
shared with the loader so the two cannot drift. Both read all four spellings
published manifests use, including the deprecated `ledmatrix_min`.

Refusal requires evidence. An undeclared floor, an unparseable version on
either side, or a core below TRUSTWORTHY_FLOOR (2.0.0) all allow the install.
That last one is deliberate and load-bearing: the v3.1.0 release reports
__version__ = "1.0.0" while nearly every published manifest floors at 2.0.0,
so a strict gate would lock those users out of the plugin store entirely --
much worse than the problem being solved. They stay unprotected until they
update the core, which is also what fixes their version string.

Verified: 782 core unit tests pass, including 25 new ones and the existing
loader-warning suite unchanged (the refactor is behavior-preserving). The
install tests drive the real install_plugin path with the download stubbed --
the allow and refuse cases differ only in the declared floor, so the refusal
is demonstrably the gate and not an earlier bail-out.

Follow-ups, deliberately not in this PR: surfacing the reason in the store UI
rather than only the log, and publishing the floor in plugins.json so the
store can refuse before downloading.

Phase B4 in docs/SPORTS_UNIFICATION.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
2026-08-02 20:33:24 -04:00
ChuckBuildsandClaude Opus 5 4f28d4eb64 fix(version): make the core version have exactly one answer
Plugin compatibility floors compare against src.__version__, so that string
has to be trustworthy. It has not been. v3.1.0 was tagged 2026-05-31 while
src/__init__.py still said "1.0.0"; the bump did not land until 2026-07-12.
Every device installed from that release reports 1.0.0, which is below the
(2, 0, 0) floor in PluginLoader._warn_if_incompatible -- so those users are
silently exempt from every plugin compatibility warning.

web_interface carried a third answer, a hardcoded "3.0.0" that nothing read
and that had drifted two majors from the core. It now re-exports the
canonical value, so it cannot disagree again.

Adds:

- test/test_version_consistency.py (enrolled in the core unit CI job):
  src.__version__ is parseable semver, matches the newest CHANGELOG heading,
  the CHANGELOG's headings are unique and descending, and web_interface
  tracks the core. src.plugin_system.__version__ is deliberately excluded --
  it versions the plugin API and moves independently.

- scripts/check_release_version.py + a release-version-check workflow that
  asserts the tag, the CHANGELOG and src.__version__ agree. Runs on pushed
  v* tags and published releases, and via workflow_dispatch so a tag can be
  checked *before* it is created:

      python scripts/check_release_version.py v3.2.0

Verified: 757 core unit tests pass including the four new ones; the script
exits 0 for v3.2.0 and non-zero for both a mismatched tag (v3.1.0) and a
non-semver one (v2.5); web_interface and web_interface.app still import.

Prerequisite for cutting v3.2.0 -- phase B4 in docs/SPORTS_UNIFICATION.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
2026-08-02 20:33:23 -04:00
11 changed files with 34 additions and 1087 deletions
@@ -1,40 +0,0 @@
name: Release version check
# A release tag, the CHANGELOG, and src.__version__ must agree. They have not
# always: v3.1.0 was tagged while src/__init__.py still said "1.0.0", which
# silently exempted every device installed from that release from plugin
# compatibility warnings. See docs/SPORTS_UNIFICATION.md (phase B4).
on:
push:
tags: ["v*"]
release:
types: [published]
# Pre-flight: run this against the tag you are about to create.
workflow_dispatch:
inputs:
tag:
description: "Tag to check (e.g. v3.2.0)"
required: true
type: string
permissions:
contents: read
jobs:
version-matches-tag:
name: Tag matches src.__version__
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
# No dependencies: the script reads src/__init__.py and CHANGELOG.md only.
- name: Assert the tag, CHANGELOG and src.__version__ agree
run: python scripts/check_release_version.py "${TAG}"
env:
TAG: ${{ inputs.tag || github.ref_name }}
+1 -4
View File
@@ -76,7 +76,4 @@ jobs:
test/test_sports_core_promotions.py \
test/test_sports_modes_promotions.py \
test/test_sports_capabilities.py \
test/test_sports_scroll.py \
test/test_version_consistency.py \
test/test_plugin_compatibility_gate.py \
test/test_install_preserves_existing.py
test/test_sports_scroll.py
-50
View File
@@ -29,21 +29,6 @@ Adoption is deliberately staged: the modules below ship here, plugins adopt them
behind guarded imports, and only then do the bundled copies go away. Nothing in
this release changes what an existing plugin loads.
**This is also the first release that *enforces* `ledmatrix_min_version`.**
Before it, the floor was advisory — the loader logged a warning and continued,
and the plugin store never compared the core version at all, so an update could
deliver a plugin that could not run. From 3.2.0 the store refuses such an
install. That matters for the sunset rule: a plugin may only delete its bundled
fallback once the cores in the field actually enforce the floor, which means
waiting for 3.2.0 to be widely installed rather than merely released. See
`docs/SPORTS_UNIFICATION.md`, phase B6.
One deliberate exception: a core reporting a version below `2.0.0` is treated as
*unknown* rather than old and is never blocked. The v3.1.0 release ships
`__version__ = "1.0.0"` (the tag was cut before the string was bumped), and
nearly every published manifest floors at `2.0.0` — so blocking on that number
would lock those users out of the plugin store entirely.
### Added
- `src/element_style.py` — per-element style resolver backing the
`x-style-elements` config-schema extension. Already consumed (behind guarded
@@ -86,37 +71,8 @@ would lock those users out of the plugin store entirely.
override point — see `docs/SPORTS_UNIFICATION.md` for where the line falls
and why.
- `src/plugin_system/compatibility.py` — the single place that answers "can this
plugin run on this core?", shared by the loader (advisory, at load time) and
the store (blocking, at install/update time) so the two cannot drift. Reads
every spelling published manifests use, including the deprecated
`versions[].ledmatrix_min`. It does **not** yet evaluate `compatible_versions`,
which is the schema-required field and can express upper bounds; closing that
is tracked in `docs/SPORTS_UNIFICATION.md` before B6.
- `scripts/check_release_version.py` and a `Release version check` workflow —
assert that a tag, the newest CHANGELOG heading and `src.__version__` agree,
on pushed `v*` tags and published releases. Runnable via `workflow_dispatch`
to check a tag *before* creating it. Added because `v3.1.0` was tagged six
weeks before `src/__init__.py` was bumped to match, which is why devices
installed from that release report `1.0.0`.
### Changed
- `src/__init__.py` bumped to **3.2.0** — the number the sunset rule keys on.
- **The plugin store refuses an incompatible install.**
`StoreManager.install_plugin` now checks the downloaded manifest's declared
floor against `src.__version__` and refuses when the plugin needs a newer
core. The check sits in `install_plugin` because `_reinstall_with_rollback`
calls it, so a refused *update* restores the version the user already had.
Refusal requires evidence: an undeclared floor, an unparseable version on
either side, or an untrustworthy core version all allow the install.
- **A failed install no longer destroys the plugin it replaced.**
`install_plugin` previously deleted the existing plugin directory before
downloading, so any later failure — a dropped connection, a malformed
manifest, or the new compatibility refusal — left the user with nothing. The
existing copy is now set aside and restored if the install fails, matching
the protection `_reinstall_with_rollback` already gave the update path.
- `web_interface.__version__` re-exports `src.__version__` instead of carrying
its own hardcoded `"3.0.0"`, which had drifted two majors from the core.
- **Live games are no longer dropped when the feed omits a game clock.**
`SportsLive._is_game_really_over` previously (in the baseball and UFC
plugin lineages) coerced a missing or non-string clock to the literal
@@ -130,12 +86,6 @@ would lock those users out of the plugin store entirely.
there means kickoff rather than expiry.
### Fixed
- **Plugin updates could hang the web request thread.** The per-plugin reinstall
locks were non-reentrant, and `_reinstall_with_rollback` holds one across its
call to `install_plugin` — which now takes the same lock to protect the
set-aside/restore above. That nesting deadlocked
`update_plugin → _reinstall_with_rollback → install_plugin`, the standard
path for every monorepo plugin update. The locks are now `RLock`s.
- `FontManager` resolves `assets/fonts` against the core install root instead
of the process working directory, so font loading works when the process
starts elsewhere (e.g. the plugin safety harness on CI).
+2 -3
View File
@@ -141,7 +141,6 @@ The system supports live, recent, and upcoming game information for multiple spo
sudo RPI_RGB_FORCE_REBUILD=1 ./first_time_install.sh
```
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and set `gpio_slowdown` to `1` or `2`.
- **1GB models (Pi 3B / 3B+) and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`.
### RGB Matrix Bonnet / HAT
@@ -315,12 +314,12 @@ curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/
```
This one-shot installer will automatically:
- Check system prerequisites (network, disk space, memory, sudo access)
- Check system prerequisites (network, disk space, sudo access)
- Install required system packages (git, python3, build tools, etc.)
- Clone or update the LEDMatrix repository
- Run the complete first-time installation script
The installation process typically takes 10-30 minutes depending on your internet connection and Pi model. Pi 3B/3B+ and other 1GB boards land at the top of that range, because the C++ library is compiled serially to stay within available memory. All errors are reported explicitly with actionable fixes.
The installation process typically takes 10-30 minutes depending on your internet connection and Pi model. All errors are reported explicitly with actionable fixes.
**Note:** The script is safe to run multiple times and will handle existing installations gracefully.
+17 -169
View File
@@ -27,7 +27,7 @@ These are independent concerns. Conflating them is what produces god classes.
| Plugin loads on a core that predates a module | Guarded import with a bundled fallback (`try: from src.X import Y / except ModuleNotFoundError: from y import Y`) |
| Plugin loads on a core that predates a *method* | Capability probing — `hasattr(SportsCore, "_detect_stale_games")` — never a version comparison. The loader's compat check is advisory-only (it logs and continues), so probing is the real protection. |
| Core changes never break a plugin's rendering | The **view-model contract**: `_extract_game_details_common` returns a dict whose `GUARANTEED_KEYS` are frozen by `test/test_skin_system.py::TestViewModelContract`. Keys may be added, never renamed or removed. |
| A plugin can drop its bundled copy safely | The **sunset rule**: its manifest must floor `ledmatrix_min_version` at the first core release shipping the module (recorded in `CHANGELOG.md`)*necessary but not sufficient*. Nothing enforces that floor today, so the copy also waits for the B6 gate below. |
| A plugin can drop its bundled copy safely | The **sunset rule**: only when its manifest floors `ledmatrix_min_version` at the first core release shipping the module (recorded in `CHANGELOG.md`). |
The core API is **additive-only**. A method the plugins call is never removed or
given a new required parameter; new behavior arrives as new methods with
@@ -201,177 +201,25 @@ legacy compatibility rather than the mechanism.
## Phases
B0B3 are merged and shipping in core 3.2.0. Everything that remains is
**rollout**, and it splits into three phases with very different risk profiles.
The original plan folded the last two together; they are separated here because
one of them is safe by construction and the other is not.
| Phase | Scope | Status | Gate |
|---|---|---|---|
| **B0** | Characterization tests, CI unit job, `element_style`, font cwd fix, CHANGELOG discipline | ✅ | — |
| **B1** | Promote the nine universal methods; convert `sports.py` → package | ✅ | Characterization suite green; no behavior change intended |
| **B2** | `CelebrationMixin` + rotation strategies as opt-in capabilities | ✅ | Non-adopters have zero new code in their MRO; strategies checked against verbatim plugin transcriptions |
| **B3** | Upstream the scroll **orchestration** layer as `src/common/sports_scroll.py`, reading `global_config['target_fps']` natively | ✅ | Content building stays per-sport |
| **B4** | Ship 3.2.0 *and* make version reporting trustworthy | ⏳ **next** | Tag, release, and `src.__version__` agree; compatibility gate merged |
| **B5** | Adoption — guarded core imports: three pilots, then the remaining six. **Bundled copies stay.** | after B4 | Per plugin: harness + goldens byte-identical, then a device soak |
| **B6** | Sunset — delete the bundled copies | **blocked** | B4's gate shipped *and* in users' hands (see below) |
### B4 — what "ship 3.2.0" actually requires
Cutting the tag is the small part. The version *number* has to become something
a floor can be trusted against, and today it is not:
- **The tag and `src.__version__` have never agreed.** `v3.1.0` was tagged
2026-05-31; `__version__` only became `"3.1.0"` on 2026-07-12 (`7f7f0d64`).
The v3.1.0 release therefore reports `__version__ = "1.0.0"`.
- **Which silences the compatibility warning entirely for that population.**
`PluginLoader._warn_if_incompatible` skips the check when the parsed core
version is below `(2, 0, 0)` — an anti-spam guard that, given the above,
matches exactly the users most likely to be behind.
- **Nothing enforces a floor anyway.** The check is advisory (it logs and
continues), and neither `StoreManager.install_plugin` nor
`StoreManager.update_plugin` compares the core version at all — `update_plugin`
compares the plugin's manifest version against the registry's
`latest_version` and nothing else.
So B4 is: tag and release 3.2.0; make the tag, the release, and `__version__`
agree, and keep them agreeing; reconsider the `< 2.0.0` skip; migrate manifests
from `ledmatrix_min` to `ledmatrix_min_version`; and add the install/update
compatibility gate that B6 depends on.
#### Two fields express compatibility, and the gate only reads one
`compatible_versions` is the canonical contract: `schema/manifest_schema.json`
**requires** it, all 42 published manifests carry it, and it holds semver
*ranges*`[">=2.0.0"]` in 41 of them, `[">=1.0.0"]` in `7-segment-clock`.
`ledmatrix_min_version` is the optional per-release floor inside `versions[]`.
The gate as merged reads only the floor. Today that is harmless: no manifest
uses an upper bound, and the two fields agree everywhere except
`7-segment-clock` (`>=1.0.0` against a `2.0.0` floor). But the fields *can*
disagree, and the range syntax the schema already permits includes upper bounds
— a plugin declaring `["2.0.0 - 2.9.9"]` means "not compatible with 3.x" and
the gate would install it on 3.2.0 regardless.
**Before B6, the gate must evaluate `compatible_versions` as well**, and the
manifest migration must reconcile the two fields rather than only renaming the
floor. Deciding which wins when they disagree is part of that work; the safe
default is the more restrictive.
(The schema also deprecates a top-level `ledmatrix_version` in favour of
`compatible_versions`. No manifest still carries it, so there is nothing to
migrate there.)
### B5 — adoption is safe by construction
A plugin adopting core imports keeps its bundled copy and reaches it through the
guarded import (see the Upgradability table above). On a core that ships the
module the plugin uses core code; on one that doesn't it falls back and behaves
exactly as it does today. There is no version of this step that breaks a user,
which is why it does not wait for B6's gate.
The hockey scroll-display pilot is **already validated**: adopted against a core
carrying 3.2.0, `scroll_display.py` went from 691 to 289 lines and all 16 harness
renders (8 sizes × 2 screens) came out byte-for-byte identical to the
pre-adoption run. That byte-comparison is the acceptance gate for every
adoption. The recipe and its two gotchas are in the plugins repo's
`docs/plugin-development/08-shared-sports-code.md`.
### B6 — why the sunset needs more than a version floor
Deleting a bundled copy removes the fallback, so the guarded import becomes a
hard dependency. On a core without the module the plugin raises
`ModuleNotFoundError` at load; `PluginManager.load_plugin` catches it, records
`PluginState.ERROR`, logs one line, and continues. Nothing crashes — the user
simply loses that scoreboard, with no visible explanation.
Verified against a `v3.1.0` worktree: `src/common/sports_scroll.py`,
`src/element_style.py` and the `src/base_classes/sports/` package are all absent
there, and the import fails with `exc.name == 'src.common.sports_scroll'`. Guard
sets must name that exact dotted path — `{"src"}` alone does not match it.
Combined with the B4 findings, a plugin that deletes its copy today reaches an
un-updated user through a normal store update, fails to load, and warns nobody.
**B6 therefore waits for B4's compatibility gate to have shipped and to have
been in users' hands long enough that the population running a core without it
is small.** The bundled copies cost disk space; deleting them early costs
scoreboards, silently. That trade is not close.
Before the first sunset, add a **compatibility regression test**. It has to
cover four cases, not one — B5's safety claim and B6's failure mode are
different propositions and only the second is obvious:
| | bundled copy present | bundled copy removed |
| Phase | Scope | Risk control |
|---|---|---|
| **pinned old core** | **loads** — this is B5's whole guarantee, that the guarded import falls back | `PluginState.ERROR`, and the recorded error names the exact missing module |
| **current core** | loads, using core code | loads, using core code |
| **B0** ✅ | Characterization tests, CI unit job, `element_style`, font cwd fix, CHANGELOG discipline | — |
| **B1** ✅ | Promote the nine universal methods; convert `sports.py` → package | Characterization suite must stay green; no behavior change intended |
| **B2** ✅ | `CelebrationMixin` + rotation strategies as opt-in capabilities | Plugins that don't opt in have zero new code in their MRO; strategies checked against verbatim plugin transcriptions |
| **B3** ✅ | Upstream the scroll **orchestration** layer as `src/common/sports_scroll.py`, reading `global_config['target_fps']` natively | Plugin copies remain until sunset; content building stays per-sport |
| **B4** | Bump to 3.2.0, record modules in CHANGELOG, migrate `ledmatrix_min``ledmatrix_min_version` | Gives plugins a version to floor on |
| **B5** ⏳ | Pilot one plugin per lineage (hockey, soccer, football) on core imports; then the remaining six; then delete bundled copies | Pilot soaks before rollout; harness + golden suites gate each |
The top-left cell is the one worth writing first: nothing in the suite currently
proves that an adopted plugin still works on a core that predates the module,
which is the entire basis for saying B5 is safe to run ahead of the gate.
**B5 is blocked on this PR merging and 3.2.0 shipping** — a plugin cannot floor
`ledmatrix_min_version` at a release that does not exist, and an unguarded
`src.common.sports_scroll` import would break every user on 3.1.0.
Assert the old-core/removed-copy case as `PluginState.ERROR` **plus the missing
module path**, not as an uncaught exception. `PluginManager.load_plugin` catches
`ModuleNotFoundError`, so nothing propagates — a test expecting a raise would
pass for the wrong reason on a core where the module is merely broken rather
than absent. "Fails loudly" is aspirational, not what the code does today: it
fails into `ERROR` state with one log line, which is precisely why B6 needs the
gate rather than trusting the failure to be noticed.
The same suite should exercise the install/update gate, since it is the other
half of the guarantee.
## What's next
In order. Each step is independently useful and independently revertible.
1. **Tag and publish v3.2.0.** The code is already on `main` (`21825cbf`).
Nothing else blocks this, and it is what makes `ledmatrix_min_version:
"3.2.0"` refer to something real.
2. **Make the version number honest.** Have the release process assert that the
tag, the GitHub release, and `src.__version__` agree — a check in CI is
cheaper than the confusion of the last two releases. Then revisit the
`< 2.0.0` skip in `_warn_if_incompatible`, which currently silences the
warning for the users who most need it.
3. **Add the compatibility gate** to `StoreManager.install_plugin` and
`.update_plugin`: refuse a plugin whose declared floor exceeds
`src.__version__`, and surface the reason in the store UI rather than only
the log. This is the single change that turns the floor from documentation
into a guarantee, and B6 depends on it.
4. **Migrate the manifests** to `ledmatrix_min_version`, and reconcile them with
`compatible_versions` (see above — that field is the required, canonical one,
and the gate does not read it yet). Currently 28 plugins spell the floor both
ways across their `versions[]` entries, 12 use only the old spelling, and 2
only the new. Scope the sweep to the nine sports plugins if a 42-plugin
version-bump wave isn't worth it — but the `compatible_versions` half has to
cover every manifest the gate can refuse, or define explicit legacy handling,
before the gate is allowed to block anything.
5. **Run B5 adoption** — hockey, soccer, football, then the remaining six.
Bundled copies stay. Byte-identical harness output per plugin, then a soak.
6. **Only then plan B6**, with the compatibility regression test described above
in CI first.
## How to keep this project healthy
Lessons this migration paid for, worth applying beyond it:
- **A version number is a promise; keep it in one place.** Three different
answers to "what version am I on" (tag, release, `__version__`) is what made
the floor untrustworthy. Assert their agreement mechanically.
- **Advisory checks protect nobody.** If a rule matters, enforce it where the
action happens — the install path, not a log line the user will never read.
If it doesn't matter enough to enforce, don't write the rule.
- **Prefer failures that are loud and early.** A plugin that dies at load with
one journal line is indistinguishable, to a user, from a plugin that was never
installed. Surface plugin health in the UI.
- **Keep the two repos' rules in sync deliberately.** The sunset rule lives in
both this file and the plugins repo's
`docs/plugin-development/08-shared-sports-code.md`. When one changes, change
the other in the same PR — drift between them is how a contributor ends up
following a rule that was superseded.
- **Measure before and after, on real hardware.** Byte-identical harness renders
and a device soak caught what unit tests could not. Reserve "it should be
fine" for things you have actually looked at.
The hockey scroll-display pilot has been **validated ahead of that gate**:
adopted against a core carrying 3.2.0, `scroll_display.py` went from 691 to 289
lines and all 16 harness renders (8 sizes × 2 screens) came out byte-for-byte
identical to the pre-adoption run. The adoption recipe and the two gotchas it
surfaced are written up in the plugins repo's
`docs/plugin-development/08-shared-sports-code.md`.
## Rules for contributors
-64
View File
@@ -82,70 +82,6 @@ python3 web_interface/start.py
## Common Issues by Category
### Installation & Build Issues
#### Step 6 fails: "Failed building wheel for rgbmatrix"
**Symptoms:**
```
note: This error originates from a subprocess, and is likely not a problem with pip.
ERROR: Failed building wheel for rgbmatrix
Failed to build rgbmatrix
✗ Failed to install rpi-rgb-led-matrix Python package
```
**Cause:**
Almost always the kernel's out-of-memory killer, not missing build tools. The
`rpi-rgb-led-matrix` library compiles roughly 45 C++ translation units, two of
them Cython-generated — a single `cc1plus` on those can peak near 800MB. The
build system defaults to running several of those at once, which exceeds RAM on
512MB and 1GB boards. Because the OOM killer writes nothing to pip's output, the
failure looks like a toolchain problem, and `sudo apt install -y
python-dev-is-python3 cmake build-essential` will report everything is already
up to date.
**How to confirm:**
```bash
dmesg -T | grep -i "out of memory" # look for "Killed process ... (cc1plus)"
free -h # total RAM and swap
```
**Fix:**
Current versions of the installer handle this automatically: they cap build
parallelism based on available RAM and add a temporary swapfile for the build,
removing it when the build finishes. If you are on an older checkout, or the
temporary swapfile could not be created, either force a serial compile:
```bash
sudo ./first_time_install.sh --build-jobs 1
```
or add permanent swap and re-run the installer, which resumes at Step 6:
```bash
sudo apt install -y dphys-swapfile
sudo sed -i 's/^#\?CONF_SWAPSIZE=.*/CONF_SWAPSIZE=2048/' /etc/dphys-swapfile
sudo sed -i 's/^#\?CONF_MAXSWAP=.*/CONF_MAXSWAP=2048/' /etc/dphys-swapfile
sudo dphys-swapfile swapoff && sudo dphys-swapfile setup && sudo dphys-swapfile swapon
sudo ./first_time_install.sh
```
`CONF_MAXSWAP` matters: it defaults to 2048 and silently clamps `CONF_SWAPSIZE`,
so setting only `CONF_SWAPSIZE` to a larger value has no effect.
**Related:**
- The installer needs roughly 3GB free on the card to place the swapfile. If
disk is tight it will say so and skip the swapfile: `sudo apt clean` first.
- `sudo bash scripts/check_system_compatibility.sh` reports RAM and disk.
- `sudo bash scripts/diagnose_dependencies.sh` dumps build-dependency state.
---
### Web Interface & Service Issues
#### Service Not Running/Starting
+8 -219
View File
@@ -152,8 +152,6 @@ ASSUME_YES=${LEDMATRIX_ASSUME_YES:-0}
SKIP_SOUND=${LEDMATRIX_SKIP_SOUND:-0}
SKIP_PERF=${LEDMATRIX_SKIP_PERF:-0}
SKIP_REBOOT_PROMPT=${LEDMATRIX_SKIP_REBOOT_PROMPT:-0}
SKIP_SWAP=${LEDMATRIX_SKIP_SWAP:-0}
BUILD_JOBS_OVERRIDE=${LEDMATRIX_BUILD_JOBS:-}
usage() {
cat <<USAGE
@@ -165,21 +163,11 @@ Options:
--skip-sound Skip sound module configuration
--skip-perf Skip performance tweaks (isolcpus/audio)
--no-reboot-prompt Do not prompt for reboot at the end
--skip-swap Never add temporary swap for the C++ build
--build-jobs N Compile the C++ library with N parallel jobs
(default: scaled to available RAM)
-h, --help Show this help message and exit
Environment variables (same effect as flags):
LEDMATRIX_ASSUME_YES=1, RPI_RGB_FORCE_REBUILD=1, LEDMATRIX_SKIP_SOUND=1,
LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1,
LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N
Low-memory devices:
On a Pi with under 2GB of RAM the C++ build is limited to fewer parallel
jobs and a temporary swapfile is added for the duration of the build, then
removed. Without this the compiler is killed by the kernel out-of-memory
killer on 512MB and 1GB models.
LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1
USAGE
}
@@ -190,38 +178,12 @@ while [ $# -gt 0 ]; do
--skip-sound) SKIP_SOUND=1 ;;
--skip-perf) SKIP_PERF=1 ;;
--no-reboot-prompt) SKIP_REBOOT_PROMPT=1 ;;
--skip-swap) SKIP_SWAP=1 ;;
--build-jobs)
shift
if [ $# -eq 0 ]; then echo "--build-jobs requires a number"; usage; exit 1; fi
BUILD_JOBS_OVERRIDE="$1"
;;
-h|--help) usage; exit 0 ;;
*) echo "Unknown option: $1"; usage; exit 1 ;;
esac
shift
done
# Low-memory build helpers (job sizing, temporary swap, OOM detection).
# Sourced rather than inlined so the sizing logic can be unit-tested; if the
# file is missing we fall back to the historical behaviour rather than failing
# the install.
LOWMEM_LIB="$PROJECT_ROOT_DIR/scripts/install/lib_lowmem.sh"
LOWMEM_AVAILABLE=0
if [ -f "$LOWMEM_LIB" ]; then
# shellcheck source=scripts/install/lib_lowmem.sh
. "$LOWMEM_LIB"
LOWMEM_AVAILABLE=1
else
echo "$LOWMEM_LIB not found; skipping low-memory build protections."
lm_remove_build_swap() { return 0; }
fi
# Remove the temporary build swapfile no matter how the script ends. Step 6
# tears it down itself; this is the backstop for the error path, since
# on_error ends in `exit` and EXIT traps still run.
trap 'lm_remove_build_swap' EXIT
# Helpers
retry() {
local attempt=1
@@ -301,144 +263,15 @@ check_disk_space() {
fi
}
# Decide how much memory Step 6's C++ build may use, and say so up front.
#
# Sets TOTAL_RAM_MB, TOTAL_SWAP_MB, BUILD_JOBS and LOW_RAM for later steps.
check_memory() {
command -v nproc >/dev/null 2>&1 && CPU_CORES=$(nproc) || CPU_CORES=1
# Validated up front rather than trusted: a non-numeric value would other-
# wise survive as far as an arithmetic test in Step 6 and fail there with a
# generic error. This must precede the fallback return below, which also
# honours the override.
if [ -n "$BUILD_JOBS_OVERRIDE" ]; then
if ! echo "$BUILD_JOBS_OVERRIDE" | grep -qE '^[1-9][0-9]*$'; then
echo "✗ Invalid build job count: '$BUILD_JOBS_OVERRIDE' (expected a positive integer)"
exit 1
fi
fi
if [ "$LOWMEM_AVAILABLE" != "1" ]; then
TOTAL_RAM_MB=0
TOTAL_SWAP_MB=0
LOW_RAM=0
BUILD_JOBS=${BUILD_JOBS_OVERRIDE:-$CPU_CORES}
return 0
fi
TOTAL_RAM_MB=$(lm_total_ram_mb)
TOTAL_SWAP_MB=$(lm_total_swap_mb)
# Test hook: exercise the low-memory path on a machine that has plenty.
if [ -n "${LEDMATRIX_FORCE_LOW_RAM:-}" ] && [ "${LEDMATRIX_FORCE_LOW_RAM}" != "0" ]; then
TOTAL_RAM_MB="${LEDMATRIX_FORCE_LOW_RAM}"
echo "⚠ LEDMATRIX_FORCE_LOW_RAM set: pretending this device has ${TOTAL_RAM_MB}MB of RAM"
fi
LOW_RAM=0
if [ "$TOTAL_RAM_MB" -gt 0 ] && [ "$TOTAL_RAM_MB" -lt 2048 ]; then
LOW_RAM=1
fi
if [ -n "$BUILD_JOBS_OVERRIDE" ]; then
BUILD_JOBS="$BUILD_JOBS_OVERRIDE"
else
BUILD_JOBS=$(lm_build_jobs "$TOTAL_RAM_MB" "$CPU_CORES")
fi
echo "System memory: ${TOTAL_RAM_MB}MB RAM, ${TOTAL_SWAP_MB}MB swap, ${CPU_CORES} core(s)"
if [ "$LOW_RAM" = "1" ]; then
echo "⚠ Low-memory device detected."
echo " The rpi-rgb-led-matrix C++ build in Step 6 will use ${BUILD_JOBS} parallel job(s)"
echo " instead of all cores, and a temporary swapfile will be added for the build"
echo " and removed afterwards. Without this the compiler is killed by the kernel"
echo " out-of-memory killer. Expect Step 6 to take 15-25 minutes."
if [ "$SKIP_SWAP" = "1" ]; then
echo " Temporary swap is disabled (--skip-swap); the build may still run out of memory."
fi
else
echo "✓ Memory sufficient for the rpi-rgb-led-matrix build (${BUILD_JOBS} parallel job(s))"
fi
}
# Compile and install the rgbmatrix Python package.
#
# CMAKE_BUILD_PARALLEL_LEVEL is the setting that actually caps the compile:
# upstream's pyproject.toml declares no [tool.scikit-build] options, so
# scikit-build-core drives Ninja through `cmake --build`, which reads this
# variable. Ninja's own default is nproc+2, i.e. six concurrent cc1plus
# processes on a 4-core Pi. MAKEFLAGS is ignored by Ninja and is set only to
# cover the Makefile-generator fallback if ninja-build is somehow absent.
#
# BUILD_TMPDIR redirects pip's build tree off tmpfs where applicable — see
# where it is computed in Step 6.
run_rgbmatrix_build() {
local jobs="$1" out="$2"
local pid elapsed=0
TMPDIR="${BUILD_TMPDIR:-${TMPDIR:-/tmp}}" \
CMAKE_BUILD_PARALLEL_LEVEL="$jobs" \
MAKEFLAGS="-j${jobs}" \
python3 -m pip install --break-system-packages . > "$out" 2>&1 &
pid=$!
# The build's output is captured to a file, so without a heartbeat a serial
# compile on a 1GB Pi looks like a 20-minute hang and invites a Ctrl-C.
#
# Polled at a short interval but reported every 30s: polling at the report
# interval instead would add most of that interval to the wall time of
# every build, including fast ones on a Pi 4/5.
while kill -0 "$pid" 2>/dev/null; do
sleep 2
elapsed=$((elapsed + 2))
if [ "$((elapsed % 30))" -eq 0 ] && kill -0 "$pid" 2>/dev/null; then
printf ' ... still compiling (%dm%02ds elapsed)\n' "$((elapsed / 60))" "$((elapsed % 60))"
fi
done
wait "$pid"
}
# Explain a failed rgbmatrix build. The kernel OOM killer writes nothing to the
# build's own output, which is why this used to be reported as a missing
# build-tools problem and sent users chasing packages they already had.
print_rgbmatrix_build_failure() {
local out="$1"
if [ "$LOWMEM_AVAILABLE" = "1" ] && lm_build_failed_on_oom "$out"; then
echo "✗ The rpi-rgb-led-matrix build was killed: the system ran out of memory."
echo " This is NOT a missing build-tools problem — the C++ compiler ran out of RAM."
echo " RAM: ${TOTAL_RAM_MB}MB Swap: $(lm_total_swap_mb)MB Parallel jobs used: ${BUILD_JOBS}"
if [ -n "${LM_SWAP_SKIP_REASON:-}" ]; then
echo " No temporary swap was added: ${LM_SWAP_SKIP_REASON}"
fi
echo ""
echo " Try one of these, then re-run this script (it resumes at Step 6):"
echo " 1. Force a single compile job:"
echo " sudo ./first_time_install.sh --build-jobs 1"
echo " 2. Add permanent swap, if the temporary swapfile could not be created:"
echo " sudo apt install -y dphys-swapfile"
echo " sudo sed -i 's/^#\\?CONF_SWAPSIZE=.*/CONF_SWAPSIZE=2048/' /etc/dphys-swapfile"
echo " sudo sed -i 's/^#\\?CONF_MAXSWAP=.*/CONF_MAXSWAP=2048/' /etc/dphys-swapfile"
echo " sudo dphys-swapfile swapoff && sudo dphys-swapfile setup && sudo dphys-swapfile swapon"
echo " 3. Free up disk space so a larger swapfile fits: sudo apt clean"
else
echo "✗ Failed to install rpi-rgb-led-matrix Python package"
echo " Ensure build tools are installed:"
echo " sudo apt install -y python-dev-is-python3 cmake build-essential"
fi
}
echo ""
echo "This script will perform the following steps:"
echo "1. Check prerequisites (network, disk, memory) and install system dependencies"
echo "1. Install system dependencies"
echo "2. Fix cache permissions"
echo "3. Fix assets directory permissions"
echo "3.1. Fix plugin directory permissions"
echo "4. Ensure configuration files exist"
echo "5. Install Python project dependencies (requirements.txt)"
echo "6. Build and install rpi-rgb-led-matrix and test import"
echo " (compiles C++; low-memory Pis get temporary swap and a serial build)"
echo "7. Install web interface dependencies"
echo "7.5. Install main LED Matrix service"
echo "8. Install web interface service"
@@ -482,16 +315,9 @@ echo "----------------------------------------"
# Pre-flight checks before APT operations
check_network
check_disk_space
check_memory
# Update package list. The one-shot installer refreshes the lists moments
# before invoking this script and exports LEDMATRIX_APT_UPDATED=1, so skip the
# duplicate refresh on that path.
if [ "${LEDMATRIX_APT_UPDATED:-0}" = "1" ]; then
echo "Package lists already refreshed by the one-shot installer; skipping apt update."
else
# Update package list
apt_update
fi
# Install required system packages
echo "Installing Python packages and dependencies..."
@@ -1076,66 +902,29 @@ else
fi
fi
# Add temporary swap on low-memory devices so the compiler survives.
CURRENT_STEP="Prepare the low-memory build environment"
if [ "$LOWMEM_AVAILABLE" = "1" ] && [ "$SKIP_SWAP" != "1" ]; then
lm_ensure_build_swap "$(lm_swap_needed_mb "$TOTAL_RAM_MB" "$TOTAL_SWAP_MB")"
elif [ "$SKIP_SWAP" = "1" ]; then
LM_SWAP_SKIP_REASON="disabled with --skip-swap"
fi
# pip builds in $TMPDIR. Debian 13 mounts /tmp as tmpfs, so the default
# would hold the entire C++ build tree in RAM — competing with the very
# compiler we are trying to keep under the memory limit.
BUILD_TMPDIR=""
if [ "$LOWMEM_AVAILABLE" = "1" ]; then
_disk_tmp=$(lm_disk_backed_tmpdir)
if [ -n "$_disk_tmp" ]; then
BUILD_TMPDIR="$_disk_tmp/ledmatrix-build"
# If this fails (a nearly-full disk being the likely cause on
# exactly the devices this targets), fall back to the default
# rather than pointing the build at a path that does not exist.
if mkdir -p "$BUILD_TMPDIR" 2>/dev/null; then
echo "Building in $BUILD_TMPDIR (TMPDIR is memory-backed; keeping the build tree on disk)"
else
echo "⚠ Could not create $BUILD_TMPDIR; falling back to the default TMPDIR"
BUILD_TMPDIR=""
fi
fi
fi
CURRENT_STEP="Build and install rpi-rgb-led-matrix"
pushd "$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master" >/dev/null
echo "Installing rpi-rgb-led-matrix Python package (scikit-build-core + cmake)..."
echo " Build deps required: python-dev-is-python3 cmake"
echo " Compiling C++ with ${BUILD_JOBS} parallel job(s)..."
if [ "$BUILD_JOBS" -le 1 ]; then
echo " Deliberately serial to stay within this device's memory — expect 15-25 minutes."
else
echo " This may take 2-5 minutes on a Pi 4/5..."
fi
echo " This compiles C++ — may take 2-5 minutes on Pi 4/5..."
BUILD_OUTPUT=$(mktemp)
BUILD_SUCCESS=false
if run_rgbmatrix_build "$BUILD_JOBS" "$BUILD_OUTPUT"; then
if python3 -m pip install --break-system-packages . > "$BUILD_OUTPUT" 2>&1; then
BUILD_SUCCESS=true
fi
cat "$BUILD_OUTPUT" >> "$LOG_FILE"
if [ "$BUILD_SUCCESS" != true ]; then
print_rgbmatrix_build_failure "$BUILD_OUTPUT"
echo "✗ Failed to install rpi-rgb-led-matrix Python package"
echo " Ensure build tools are installed:"
echo " sudo apt install -y python-dev-is-python3 cmake build-essential"
echo ""
echo "-- Last 50 lines of build output --"
tail -n 50 "$BUILD_OUTPUT"
rm -f "$BUILD_OUTPUT"
if [ -n "$BUILD_TMPDIR" ]; then rm -rf "$BUILD_TMPDIR"; fi
popd >/dev/null
lm_remove_build_swap
exit 1
fi
rm -f "$BUILD_OUTPUT"
if [ -n "$BUILD_TMPDIR" ]; then rm -rf "$BUILD_TMPDIR"; fi
popd >/dev/null
# Hand the memory back well before Step 14's reboot.
lm_remove_build_swap
else
echo "✗ rpi-rgb-led-matrix-master directory not found at $PROJECT_ROOT_DIR"
echo "Failed to initialize submodule or clone repository"
+2 -6
View File
@@ -156,13 +156,9 @@ echo ""
echo "6. Check disk space - building packages requires temporary space"
echo " df -h"
echo ""
echo "7. For slow builds or out-of-memory kills, increase swap space."
echo " first_time_install.sh already adds temporary swap on low-memory devices;"
echo " this makes it permanent. Set CONF_MAXSWAP too - it defaults to 2048 and"
echo " silently clamps CONF_SWAPSIZE, so raising CONF_SWAPSIZE alone does nothing."
echo "7. For slow builds, increase swap space:"
echo " sudo dphys-swapfile swapoff"
echo " sudo sed -i 's/^#\\?CONF_SWAPSIZE=.*/CONF_SWAPSIZE=2048/' /etc/dphys-swapfile"
echo " sudo sed -i 's/^#\\?CONF_MAXSWAP=.*/CONF_MAXSWAP=2048/' /etc/dphys-swapfile"
echo " sudo nano /etc/dphys-swapfile # Set CONF_SWAPSIZE=2048"
echo " sudo dphys-swapfile setup"
echo " sudo dphys-swapfile swapon"
echo ""
-283
View File
@@ -1,283 +0,0 @@
#!/bin/bash
#
# Low-memory build helpers for the LED Matrix installer.
#
# Sourced by first_time_install.sh. These live in a separate, sourceable file
# so the pure sizing/detection functions can be unit-tested
# (test/test_install_lowmem.py); first_time_install.sh itself is not sourceable
# because it self-elevates and runs top to bottom.
#
# Why this exists: the rgbmatrix build compiles ~45 C++ translation units, two
# of them Cython-generated (a single cc1plus on those peaks around 400-800MB at
# -O3). Upstream's pyproject.toml sets no [tool.scikit-build] options, so
# scikit-build-core uses Ninja at its default of nproc+2 jobs -- six concurrent
# compiles on a 4-core Pi. On a 512MB-1GB Pi the OOM killer reaps cc1plus and
# pip reports only "Failed building wheel for rgbmatrix".
#
# The caller runs under `set -Eeuo pipefail` with an ERR trap, and these are
# invoked from the middle of numbered steps, so nothing here may call exit and
# the swap helpers must always return 0.
# Overridable so tests can point at fixture files instead of /proc.
LM_MEMINFO="${LM_MEMINFO:-/proc/meminfo}"
LM_SWAPS="${LM_SWAPS:-/proc/swaps}"
# Temporary swapfile created for the build and removed afterwards. Deliberately
# never added to /etc/fstab: a malformed fstab can leave a novice with an
# unbootable Pi, and this swap only needs to outlive the compile.
LM_SWAPFILE="${LM_SWAPFILE:-/var/swap.ledmatrix-install}"
# Bring RAM + real swap up to this much before compiling, capped per swapfile.
LM_SWAP_TARGET_MB="${LM_SWAP_TARGET_MB:-3072}"
LM_SWAP_MAX_MB="${LM_SWAP_MAX_MB:-2048}"
# Worst-case cc1plus footprint on the Cython translation unit, used to size
# build parallelism against available RAM.
LM_MB_PER_JOB="${LM_MB_PER_JOB:-768}"
# Set to 1 once swap is live, so lm_remove_build_swap (wired up as an EXIT
# trap) knows whether there is anything to undo.
LM_TEMP_SWAP_ACTIVE=0
# Human-readable reason no swapfile was created, quoted back in the failure
# message so a user who still OOMs is told why the safety net was absent.
LM_SWAP_SKIP_REASON=""
# ---------------------------------------------------------------------------
# Pure helpers (no side effects; unit-tested)
# ---------------------------------------------------------------------------
# Total physical RAM in MB, or 0 if it cannot be determined.
lm_total_ram_mb() {
# Defaults are re-resolved here as well as at source time so the function
# stays safe under the installer's `set -u`.
awk '/^MemTotal:/ {printf "%d\n", $2 / 1024; found = 1; exit} END {if (!found) print 0}' \
"${LM_MEMINFO:-/proc/meminfo}" 2>/dev/null || echo 0
}
# Total swap in MB, EXCLUDING zram devices.
#
# zram swap is compressed RAM: it consumes the very resource that is already
# exhausted and does nothing for a build OOM. Counting it would let a
# zram-enabled image decide it has enough swap and then fail exactly as before.
lm_total_swap_mb() {
awk 'NR > 1 && $1 !~ /^\/dev\/zram/ {total += $3} END {printf "%d\n", total / 1024}' \
"${LM_SWAPS:-/proc/swaps}" 2>/dev/null || echo 0
}
# lm_build_jobs <ram_mb> <cores> -> max(1, min(cores, ram_mb / LM_MB_PER_JOB))
#
# Computed from RAM alone and never RAM+swap: handing out extra jobs because
# swap exists just guarantees SD-card thrash, which is far slower than
# compiling serially.
lm_build_jobs() {
local ram_mb="${1:-0}" cores="${2:-1}" jobs
local per_job="${LM_MB_PER_JOB:-768}"
if [ "$cores" -lt 1 ]; then
cores=1
fi
jobs=$(( ram_mb / per_job ))
if [ "$jobs" -lt 1 ]; then
jobs=1
fi
if [ "$jobs" -gt "$cores" ]; then
jobs="$cores"
fi
echo "$jobs"
}
# lm_swap_needed_mb <ram_mb> <existing_swap_mb> -> swapfile size in MB, or 0.
#
# Brings RAM + real swap up to LM_SWAP_TARGET_MB, capped at LM_SWAP_MAX_MB and
# rounded up to a 256MB multiple. Machines with enough memory get 0 and are
# left completely untouched.
lm_swap_needed_mb() {
local ram_mb="${1:-0}" swap_mb="${2:-0}" needed
local target="${LM_SWAP_TARGET_MB:-3072}" max="${LM_SWAP_MAX_MB:-2048}"
needed=$(( target - ram_mb - swap_mb ))
if [ "$needed" -le 0 ]; then
echo 0
return 0
fi
if [ "$needed" -gt "$max" ]; then
needed="$max"
fi
echo $(( ( (needed + 255) / 256 ) * 256 ))
}
# lm_build_failed_on_oom <build_output_file> -> 0 if the build was OOM-killed.
#
# Two independent evidence sources, because neither alone is reliable: the
# compiler sometimes reports its own allocation failure, but when the kernel
# OOM killer fires it writes nothing to the build's stdout. That silence is
# exactly why the old handler misdiagnosed this as missing build tools.
lm_build_failed_on_oom() {
local build_output="${1:-}" kernel_log=""
if [ -n "$build_output" ] && [ -f "$build_output" ]; then
if grep -qiE 'cc1plus: out of memory|virtual memory exhausted|Cannot allocate memory|MemoryError|fatal error: Killed signal terminated program|signal 9' \
"$build_output"; then
return 0
fi
fi
# LM_KERNEL_LOG_FILE lets tests supply a fixture instead of the real kernel
# ring buffer, which on a shared CI machine may hold unrelated OOM events.
if [ -n "${LM_KERNEL_LOG_FILE:-}" ]; then
if [ -f "$LM_KERNEL_LOG_FILE" ]; then
kernel_log=$(cat "$LM_KERNEL_LOG_FILE" 2>/dev/null || true)
fi
elif command -v dmesg >/dev/null 2>&1; then
kernel_log=$(dmesg -T 2>/dev/null || dmesg 2>/dev/null || true)
fi
if [ -z "$kernel_log" ] && [ -z "${LM_KERNEL_LOG_FILE:-}" ] && command -v journalctl >/dev/null 2>&1; then
kernel_log=$(journalctl -k --since "30 min ago" --no-pager 2>/dev/null || true)
fi
if [ -n "$kernel_log" ]; then
if printf '%s\n' "$kernel_log" | tail -n 300 | \
grep -qiE 'Out of memory: Kill|oom_kill|oom-kill|Killed process'; then
return 0
fi
fi
return 1
}
# lm_disk_backed_tmpdir [candidate] -> a disk-backed temp dir, or nothing.
#
# pip builds in $TMPDIR. Debian 13 mounts /tmp as tmpfs, so the default puts the
# whole C++ build tree in RAM, competing with the compiler we are already trying
# to keep under the limit. Prints a replacement only when the current TMPDIR is
# memory-backed and the candidate is not; otherwise prints nothing and the
# caller keeps its default.
lm_disk_backed_tmpdir() {
local candidate="${1:-/var/tmp}"
local current="${TMPDIR:-/tmp}"
local current_fs="" candidate_fs=""
current_fs=$(lm_fstype_of "$current")
case "$current_fs" in
tmpfs|ramfs) ;;
*) return 0 ;;
esac
candidate_fs=$(lm_fstype_of "$candidate")
case "$candidate_fs" in
tmpfs|ramfs|"") return 0 ;;
esac
echo "$candidate"
}
# Filesystem type backing a path, or empty if it cannot be determined.
lm_fstype_of() {
local path="${1:-/}"
if command -v findmnt >/dev/null 2>&1; then
findmnt -no FSTYPE --target "$path" 2>/dev/null | head -n 1
return 0
fi
if command -v stat >/dev/null 2>&1; then
stat -f -c %T "$path" 2>/dev/null | head -n 1
return 0
fi
return 0
}
# ---------------------------------------------------------------------------
# Swap management (requires root; not unit-tested)
# ---------------------------------------------------------------------------
# lm_ensure_build_swap <needed_mb>
#
# Always returns 0. On any refusal it sets LM_SWAP_SKIP_REASON and leaves the
# system untouched -- swap is a safety net for the build, never a precondition.
lm_ensure_build_swap() {
local needed_mb="${1:-0}"
local swap_dir free_mb budget
LM_SWAP_SKIP_REASON=""
if [ "$needed_mb" -le 0 ]; then
LM_SWAP_SKIP_REASON="not needed (RAM and existing swap are sufficient)"
return 0
fi
if ! command -v mkswap >/dev/null 2>&1 || ! command -v swapon >/dev/null 2>&1; then
LM_SWAP_SKIP_REASON="mkswap/swapon are not available on this system"
echo "⚠ Cannot add build swap: $LM_SWAP_SKIP_REASON"
return 0
fi
# Clear a stale swapfile left by a run that was killed before its cleanup
# ran, so this is safe to call repeatedly.
if [ -e "$LM_SWAPFILE" ]; then
echo "Removing a leftover swapfile from a previous run: $LM_SWAPFILE"
swapoff "$LM_SWAPFILE" >/dev/null 2>&1 || true
rm -f "$LM_SWAPFILE" || true
fi
# Keep a working margin for the build tree itself; never eat the last GB.
swap_dir=$(dirname "$LM_SWAPFILE")
free_mb=$(df -m "$swap_dir" 2>/dev/null | awk 'NR==2{print $4}')
free_mb=${free_mb:-0}
budget=$(( free_mb - 1024 ))
if [ "$budget" -lt 256 ]; then
LM_SWAP_SKIP_REASON="only ${free_mb}MB free on ${swap_dir}, need about $(( needed_mb + 1024 ))MB"
echo "⚠ Skipping the build swapfile: $LM_SWAP_SKIP_REASON"
return 0
fi
if [ "$needed_mb" -gt "$budget" ]; then
echo "⚠ Trimming the build swapfile from ${needed_mb}MB to leave 1GB free on ${swap_dir}"
needed_mb=$(( ( budget / 256 ) * 256 ))
fi
echo "Adding a temporary ${needed_mb}MB swapfile for the build: $LM_SWAPFILE"
echo " This is removed automatically once the build finishes."
# fallocate can produce a sparse file that mkswap rejects, and is not
# supported on every filesystem; dd always yields a usable file.
if ! fallocate -l "${needed_mb}M" "$LM_SWAPFILE" 2>/dev/null; then
if ! dd if=/dev/zero of="$LM_SWAPFILE" bs=1M count="$needed_mb" status=none 2>/dev/null; then
LM_SWAP_SKIP_REASON="could not allocate ${needed_mb}MB at $LM_SWAPFILE"
echo "$LM_SWAP_SKIP_REASON"
rm -f "$LM_SWAPFILE" || true
return 0
fi
fi
chmod 600 "$LM_SWAPFILE" || true
if ! mkswap "$LM_SWAPFILE" >/dev/null 2>&1; then
LM_SWAP_SKIP_REASON="mkswap failed on $LM_SWAPFILE"
echo "$LM_SWAP_SKIP_REASON"
rm -f "$LM_SWAPFILE" || true
return 0
fi
if ! swapon "$LM_SWAPFILE" >/dev/null 2>&1; then
LM_SWAP_SKIP_REASON="swapon failed on $LM_SWAPFILE"
echo "$LM_SWAP_SKIP_REASON"
rm -f "$LM_SWAPFILE" || true
return 0
fi
LM_TEMP_SWAP_ACTIVE=1
echo "✓ Temporary build swap active (${needed_mb}MB; total swap is now $(lm_total_swap_mb)MB)"
return 0
}
# Remove the temporary swapfile. Safe to call unconditionally and repeatedly.
#
# Wired up as an EXIT trap, so it must never return non-zero -- a failing trap
# would surface as a spurious installer error.
lm_remove_build_swap() {
if [ "${LM_TEMP_SWAP_ACTIVE:-0}" != "1" ]; then
return 0
fi
LM_TEMP_SWAP_ACTIVE=0
echo "Removing the temporary build swapfile: $LM_SWAPFILE"
swapoff "$LM_SWAPFILE" >/dev/null 2>&1 || true
rm -f "$LM_SWAPFILE" || true
return 0
}
+3 -39
View File
@@ -145,34 +145,6 @@ check_disk_space() {
fi
}
# Report available memory so the user knows what to expect before the wait.
#
# Informational only — first_time_install.sh does the real work of capping
# build parallelism and adding temporary swap. Never fatal: a low-RAM Pi is
# supported, it is just slower.
check_memory() {
CURRENT_STEP="Memory check"
if [ ! -r /proc/meminfo ]; then
print_warning "Cannot read /proc/meminfo, skipping memory check"
return 0
fi
TOTAL_RAM_MB=$(awk '/^MemTotal:/ {printf "%d\n", $2 / 1024; exit}' /proc/meminfo 2>/dev/null || echo 0)
TOTAL_RAM_MB=${TOTAL_RAM_MB:-0}
if [ "$TOTAL_RAM_MB" -eq 0 ]; then
print_warning "Could not determine system memory, continuing"
elif [ "$TOTAL_RAM_MB" -lt 2048 ]; then
print_warning "Low memory: ${TOTAL_RAM_MB}MB RAM"
echo " The rpi-rgb-led-matrix C++ build needs more memory than this Pi has."
echo " The installer will compile with fewer parallel jobs and add a temporary"
echo " swapfile for the build, removing it afterwards. That step will take"
echo " 15-25 minutes rather than the usual 2-5."
else
print_success "Memory sufficient: ${TOTAL_RAM_MB}MB RAM"
fi
}
# Ensure sudo access
check_sudo() {
CURRENT_STEP="Sudo access check"
@@ -232,7 +204,7 @@ main() {
print_step "LED Matrix One-Shot Installation"
echo "This script will:"
echo " 1. Check prerequisites (network, disk space, memory, sudo)"
echo " 1. Check prerequisites (network, disk space, sudo)"
echo " 2. Install system dependencies (git, python3, build tools)"
echo " 3. Clone the LEDMatrix repository"
echo " 4. Run the first-time installation script"
@@ -241,7 +213,6 @@ main() {
# Check prerequisites
check_network
check_disk_space
check_memory
check_sudo
# Note: /tmp permissions are checked and fixed inline before running first_time_install.sh
# (only if actually wrong, not preemptively)
@@ -257,14 +228,12 @@ main() {
exit 1
fi
# Update package list first. first_time_install.sh is told the lists are
# already fresh so it does not repeat this a minute later.
# Update package list first
if [ "$EUID" -eq 0 ]; then
retry apt-get update -qq
else
retry sudo apt-get update -qq
fi
export LEDMATRIX_APT_UPDATED=1
# Install git and curl (needed for cloning and the script itself)
if ! command -v git >/dev/null 2>&1 || ! command -v curl >/dev/null 2>&1; then
@@ -403,12 +372,7 @@ main() {
# Pass both -y flag AND environment variable for non-interactive mode
# This ensures it works even if the script re-executes itself with sudo
# Also ensure stdin is properly handled for non-interactive mode
# LEDMATRIX_APT_UPDATED is passed explicitly rather than relying on
# -E: a sudoers env_reset/env_keep policy can strip exported variables,
# which would silently reinstate the duplicate apt update.
sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 \
LEDMATRIX_APT_UPDATED="${LEDMATRIX_APT_UPDATED:-0}" \
bash ./first_time_install.sh -y </dev/null
sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 bash ./first_time_install.sh -y </dev/null
fi
INSTALL_EXIT_CODE=$?
trap 'on_error $LINENO' ERR # Re-enable ERR trap
-209
View File
@@ -1,209 +0,0 @@
"""
Tests for scripts/install/lib_lowmem.sh, the installer's low-memory helpers.
Background: the rgbmatrix build compiles ~45 C++ translation units, two of them
Cython-generated. Upstream's pyproject.toml sets no [tool.scikit-build] options,
so scikit-build-core drives Ninja at its default of nproc+2 jobs -- six
concurrent cc1plus on a 4-core Pi. On 512MB and 1GB models the OOM killer reaps
the compiler and pip reports only "Failed building wheel for rgbmatrix", which
the installer used to misreport as a missing-build-tools problem.
These cover the pure sizing/detection functions. The swap-management functions
need root and mutate the system, so they are exercised manually instead.
"""
import subprocess
from pathlib import Path
import pytest
LIB = Path(__file__).resolve().parent.parent / "scripts" / "install" / "lib_lowmem.sh"
def run_lib(snippet: str, env: dict | None = None) -> subprocess.CompletedProcess:
"""Source the helper library and run a snippet against it."""
script = f". {LIB}\n{snippet}"
return subprocess.run(
["bash", "-c", script],
capture_output=True,
text=True,
env={"PATH": "/usr/bin:/bin:/usr/sbin:/sbin", **(env or {})},
)
def call(fn: str, *args: object, env: dict | None = None) -> str:
joined = " ".join(str(a) for a in args)
result = run_lib(f"{fn} {joined}", env=env)
assert result.returncode == 0, f"{fn} failed: {result.stderr}"
return result.stdout.strip()
class TestLibraryLoads:
def test_library_exists_and_is_syntactically_valid(self):
assert LIB.is_file(), f"{LIB} is missing"
result = subprocess.run(["bash", "-n", str(LIB)], capture_output=True, text=True)
assert result.returncode == 0, result.stderr
def test_sourcing_is_safe_under_strict_mode(self):
# first_time_install.sh runs under `set -Eeuo pipefail` with an ERR
# trap, so sourcing must not trip either.
result = run_lib("set -Eeuo pipefail\ntrap 'exit 99' ERR\necho ok")
assert result.returncode == 0, result.stderr
assert "ok" in result.stdout
class TestBuildJobs:
@pytest.mark.parametrize(
"ram_mb,cores,expected",
[
(512, 4, 1), # Pi Zero 2 W - must serialize
(1024, 4, 1), # Pi 3B/3B+ - the device from the bug report
(2048, 4, 2),
(4096, 4, 4), # core-capped
(8192, 4, 4), # core-capped
(2048, 1, 1), # single-core machine
],
)
def test_jobs_scale_with_ram_and_cap_at_cores(self, ram_mb, cores, expected):
assert call("lm_build_jobs", ram_mb, cores) == str(expected)
def test_never_returns_zero_jobs(self):
assert call("lm_build_jobs", 0, 4) == "1"
def test_treats_zero_cores_as_one(self):
assert call("lm_build_jobs", 8192, 0) == "1"
class TestSwapSizing:
@pytest.mark.parametrize(
"ram_mb,existing_swap_mb,expected",
[
(512, 0, 2048), # capped at LM_SWAP_MAX_MB
(1024, 0, 2048), # matches the workaround the reporter found
(2048, 0, 1024),
(2048, 1024, 0), # existing swap already covers it
(4096, 0, 0), # untouched on machines that already work
(8192, 0, 0),
],
)
def test_swap_target_scales_with_ram(self, ram_mb, existing_swap_mb, expected):
assert call("lm_swap_needed_mb", ram_mb, existing_swap_mb) == str(expected)
def test_result_is_a_multiple_of_256mb(self):
# 3072 - 900 = 2172, which must round up rather than produce an odd size.
assert int(call("lm_swap_needed_mb", 900, 0)) % 256 == 0
class TestSwapDetection:
def test_zram_swap_is_excluded(self, tmp_path):
# zram swap is compressed RAM: counting it would let a zram-enabled
# image skip provisioning and then OOM exactly as before.
swaps = tmp_path / "swaps"
swaps.write_text(
"Filename\t\t\t\tType\t\tSize\t\tUsed\t\tPriority\n"
"/dev/zram0 partition\t1048572\t\t0\t\t100\n"
"/var/swap file\t\t524284\t\t0\t\t-2\n"
)
assert call("lm_total_swap_mb", env={"LM_SWAPS": str(swaps)}) == "511"
def test_zram_only_system_reports_no_usable_swap(self, tmp_path):
swaps = tmp_path / "swaps"
swaps.write_text(
"Filename\t\t\t\tType\t\tSize\t\tUsed\t\tPriority\n"
"/dev/zram0 partition\t1048572\t\t0\t\t100\n"
)
assert call("lm_total_swap_mb", env={"LM_SWAPS": str(swaps)}) == "0"
def test_ram_is_read_from_meminfo(self, tmp_path):
meminfo = tmp_path / "meminfo"
# A real Pi 3B+ reports this; 948204/1024 truncates to 925.
meminfo.write_text("MemTotal: 948204 kB\nMemFree: 123456 kB\n")
assert call("lm_total_ram_mb", env={"LM_MEMINFO": str(meminfo)}) == "925"
def test_missing_files_report_zero_rather_than_failing(self, tmp_path):
missing = str(tmp_path / "nope")
assert call("lm_total_ram_mb", env={"LM_MEMINFO": missing}) == "0"
assert call("lm_total_swap_mb", env={"LM_SWAPS": missing}) == "0"
class TestOomDetection:
"""The regression tests for the misdiagnosis in the bug report."""
def _check(self, tmp_path, build_log: str, kernel_log: str = "") -> bool:
build_file = tmp_path / "build.log"
build_file.write_text(build_log)
kernel_file = tmp_path / "kernel.log"
kernel_file.write_text(kernel_log)
result = run_lib(
f"lm_build_failed_on_oom {build_file}",
env={"LM_KERNEL_LOG_FILE": str(kernel_file)},
)
return result.returncode == 0
@pytest.mark.parametrize(
"line",
[
"c++: fatal error: Killed signal terminated program cc1plus",
"cc1plus: out of memory allocating 65536 bytes",
"virtual memory exhausted: Cannot allocate memory",
"error: command '/usr/bin/c++' died with signal 9",
],
)
def test_detects_compiler_reported_memory_failures(self, tmp_path, line):
log = f"[15/45] Building CXX object core.cpp.o\n{line}\nninja: build stopped.\n"
assert self._check(tmp_path, log) is True
def test_detects_oom_visible_only_in_the_kernel_log(self, tmp_path):
# The OOM killer writes nothing to the build's stdout. This silence is
# precisely why the old handler blamed missing build tools.
build_log = (
"[15/45] Building CXX object core.cpp.o\n"
"ninja: build stopped: subcommand failed.\n"
"ERROR: Failed building wheel for rgbmatrix\n"
)
kernel_log = (
"[12345.6] Out of memory: Killed process 4242 (cc1plus) "
"total-vm:812345kB, anon-rss:764000kB\n"
)
assert self._check(tmp_path, build_log, kernel_log) is True
def test_does_not_flag_a_genuine_missing_build_tool(self, tmp_path):
build_log = (
"CMake Error at CMakeLists.txt:12 (find_package):\n"
" Could NOT find Python (missing: Development.Module)\n"
"fatal error: Python.h: No such file or directory\n"
"ERROR: Failed building wheel for rgbmatrix\n"
)
assert self._check(tmp_path, build_log) is False
def test_does_not_flag_a_network_failure(self, tmp_path):
build_log = (
"WARNING: Retrying after connection broken by 'NewConnectionError'\n"
"ERROR: Could not install packages due to an OSError\n"
)
assert self._check(tmp_path, build_log) is False
def test_handles_a_missing_build_log(self, tmp_path):
kernel_file = tmp_path / "kernel.log"
kernel_file.write_text("")
result = run_lib(
f"lm_build_failed_on_oom {tmp_path / 'absent.log'}",
env={"LM_KERNEL_LOG_FILE": str(kernel_file)},
)
assert result.returncode == 1
class TestDiskBackedTmpdir:
def test_returns_nothing_when_tmpdir_is_already_disk_backed(self, tmp_path):
# tmp_path is on the regular filesystem, so the default must be kept.
assert call("lm_disk_backed_tmpdir", env={"TMPDIR": str(tmp_path)}) == ""
def test_redirects_away_from_a_memory_backed_tmpdir(self):
# Debian 13 mounts /tmp as tmpfs, which would otherwise hold the whole
# C++ build tree in RAM alongside the compiler.
shm = Path("/dev/shm")
if not shm.is_dir():
pytest.skip("/dev/shm not available")
result = run_lib("lm_disk_backed_tmpdir", env={"TMPDIR": str(shm)})
assert result.returncode == 0
assert result.stdout.strip() in ("", "/var/tmp")