Test suite overhaul + fixes for the three bugs it uncovered (#441)

* ci: run the whole test tree and make the plugin-safety job assert something real

The unit-tests CI job ran an explicit 24-file allowlist that had rotted:
63 of 90 test files (display, vegas, store manager, web API, web_interface)
never ran on a PR. The job now runs all of test/ (minus test/plugins, which
the plugin-safety job owns) so new test files are enrolled by default and
any exclusion needs a visible, commented --ignore.

The plugin-safety job was a green no-op: plugins/ is empty in CI, so every
test skipped with 'Manifest not found'. It now renders a bundled
deterministic fixture plugin (test/fixtures/plugins/ci-fixture-plugin,
golden images included for all 8 default sizes) via LEDMATRIX_PLUGINS_DIR,
and sets LEDMATRIX_REQUIRE_PLUGINS=1 so discovering zero plugins fails
loudly instead of skipping green. The per-plugin suites document that they
target dev machines with real plugins installed.

Coverage is now measured and enforced in exactly one place — the CI
unit-tests step (--cov=src --cov=web_interface --cov-fail-under=45, from a
measured 47% baseline). pytest.ini previously declared --cov-fail-under=30
but CI always passed --no-cov, so the gate had never run anywhere; local
pytest is now coverage-free and fast.

Enabling the 63 unenrolled files surfaced three cases of test rot, fixed
here: test_display_controller_vegas_tick.py could not collect without the
hardware rgbmatrix module (now uses the emulator convention), the
state-reconciliation unrecoverable-cache tests broke when production added
the is_plugin_uninstalled tombstone check (bare Mock returned truthy),
and test_get_system_status assumed the optional psutil dependency
(now installed via requirements-test.txt and guarded by importorskip).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: replace can't-fail tests with real assertions

test_font_manager.py was 5 of 6 tests shaped as 'try: call(); assert True /
except: assert True' — running in CI while unable to fail on any
regression. Rewritten against the real FontManager API and the bundled
assets/fonts: returned font types, cache-hit identity, distinct entries per
size, default-font fallback for unknown families and corrupt files
(recorded in failed_loads), BDF native-size reading, text measurement, and
cache lifecycle.

test_display_manager.py's test_draw_text ended in 'assert True'; it now
renders onto a known-black canvas and asserts pixels were actually lit —
which required un-breaking the fixture's freetype MagicMock so draw_text's
isinstance check doesn't silently swallow the draw.

test_display_controller.py carried a permanently-skipped test whose skip
reason already declared it redundant; deleted.

Both display test files now set EMULATOR=true before importing
display_manager (the same convention as test_display_dirty_tracking.py) so
they collect standalone instead of depending on which test module imports
display_manager first.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: cover the untested fragile logic (compatibility gate, secrets, config merges, durations, skin cards)

New unit tests for pure or filesystem-only logic that previously had zero
direct coverage:

- test_compatibility.py: the semver install gate (parse_semver suffix
  handling, every range operator, TRUSTWORTHY_FLOOR behavior for cores
  reporting untrustworthy versions, 'more restrictive wins', and the
  malformed-manifest shapes that used to raise).
- test/web_interface/test_secret_helpers.py: the canonical x-secret
  helpers — find/separate/mask/remove, array-item secrets, no input
  mutation, and a separate->recombine round-trip.
- test/web_interface/test_api_v3_helpers.py: the module-level helpers
  behind the plugin config save endpoint (_is_plugin_update_available,
  _coerce_to_bool including the int==1 quirk, deep_merge including its
  shared-subtree shallowness, _parse_form_value, dotted-key-aware
  _get_schema_property/_set_nested_value).
- test_base_plugin_duration.py: get_display_duration's full coercion
  ladder (instance attr -> config -> 15.0), including the bool-is-int
  quirk where display_duration=True means one second.
- test_config_manager_secrets.py: the secrets round-trip — deep-merge on
  load, strip on save, group pruning, the load fast path — and two
  characterized sharp edges marked SUSPECTED BUG: an unreadable secrets
  file at save time writes secrets into config.json in plaintext, and a
  same-mtime-same-size content swap is served stale.
- test_schema_manager_merge.py: merge_with_defaults branch behavior (None
  replacement vs falsey preservation, dict-vs-scalar mismatches, arrays
  replaced wholesale, defaults never mutated).
- test_skin_system.py (extended): render_skin_card shares _render_game's
  3-strike counter but never resets it on success — the asymmetry is
  pinned in both directions, along with card fallthrough and the disable
  interaction between the two paths.

Suspected bugs are characterized, not fixed — each carries a comment so a
future behavior change is deliberate rather than accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: add drift guards for cross-file contracts

Three guard suites that pin contracts spanning multiple files, where one
side changing unilaterally breaks the other silently:

- test_version_comparison_consistency.py: the repo's four version
  comparators (compatibility.parse_semver, api_v3's packaging-based
  _is_plugin_update_available, store_manager update_plugin's raw string
  equality, skin_runtime._major) answer differently on the same inputs.
  A table pins each one's verdict; update_plugin is driven through its
  real code path to show the SUSPECTED BUGs: 'v1.2.0' vs '1.2.0'
  triggers a full reinstall the UI calls unnecessary, and a locally-ahead
  plugin gets downgraded. A pairwise-ordering check keeps parse_semver
  agreeing with packaging on plain X.Y.Z.
- test/web_interface/test_secret_separation_parity.py: api_v3.py carries
  three inline copies of find_secret_fields/separate_secrets that lack
  the canonical module's array-item support. The copy count is asserted
  exact (it may only go down; new copies must import
  src/web_interface/secret_helpers), the missing-array-support gap is
  asserted so it can't grow silently, and the canonical behavior that
  migration will adopt is documented executably.
- test_discovery_path_contract.py: the three 'where is plugin X'
  resolvers (PluginManager discovery, StoreManager._find_plugin_path,
  SchemaManager.get_schema_path) agree on the configured directory, and
  their divergent fallback chains are characterized. Also pins the
  .standalone-backup- naming contract shared by store rollback and
  discovery, and _resolve_skin_target's path-traversal rejection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: address review feedback — fixture lifecycle, test names, ClassVar

- ci-fixture-plugin: call display_manager.clear() before rendering (per
  plugin guidelines — the fixture should model a well-behaved plugin),
  add a class docstring, and document why Pillow is deliberately not
  pinned in its requirements.txt (core dependency; harness installs
  nothing).
- Rename two tests whose names contradicted their assertions:
  test_unparseable_core_version_is_compatible ->
  test_unparseable_core_with_high_floor_is_blocked, and
  test_unreadable_secrets_file... -> test_corrupt_secrets_file...
- Annotate TestGetSchemaProperty.SCHEMA as ClassVar (RUF012).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* ci: allow manual test.yml runs via workflow_dispatch

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: unify version comparison, refuse secret-leaking saves, reset skin strikes on card success

Fixes the three suspected bugs this PR's characterization tests pinned,
flipping those tests to assert the corrected behavior:

- plugins/store: ONE shared update comparator. New
  compatibility.is_update_available() (PEP 440 via packaging) is now used
  by both the web UI's update badge (api_v3._is_plugin_update_available
  is a thin alias) and store_manager.update_plugin's reinstall decision.
  Previously update_plugin used raw string equality: 'v1.2.0' vs '1.2.0'
  triggered a full reinstall the UI called unnecessary, and a locally-
  ahead plugin (2.0.0 installed, registry 1.9.0) was silently DOWNGRADED.
  Now equivalent spellings skip the reinstall and locally-ahead versions
  are never downgraded; unparseable versions still reconcile by
  reinstalling from the registry.

- config: save_config and save_config_atomic now refuse (ConfigError)
  when config_secrets.json exists but cannot be loaded. Both previously
  proceeded without stripping, writing the merged secrets into
  config.json in plaintext. The shared _load_secrets_for_save() helper
  raises with an actionable message instead; a missing secrets file is
  still fine (nothing to strip), and _migrate_config's catch-all keeps
  boot resilient.

- skins: render_skin_card resets _skin_failures on both success paths
  (vegas card returned, or mode renderer handled), mirroring
  _render_game. Transient card failures no longer accumulate across a
  session until they permanently disable a working skin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: harden shared comparator edges from review

- is_update_available: reject truthy non-string versions (a malformed
  manifest can carry a number; packaging raises TypeError on those) by
  surfacing the mismatch instead of raising.
- store_manager.update_plugin: drop the truthiness gate around the
  comparator so a missing version on either side follows the shared
  'no update' verdict, keeping the store consistent with the UI badge;
  a missing manifest still uses the reinstall recovery path.
- config_manager._load_secrets_for_save: catch only expected read/parse
  failures (OSError/ValueError/RecursionError) so implementation bugs
  propagate as themselves, and log with traceback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

---------

Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-08-07 10:17:30 -04:00
committed by GitHub
co-authored by Claude Fable 5
parent d9683e28be
commit d6c5f97c13
43 changed files with 2147 additions and 184 deletions
+120 -76
View File
@@ -1,82 +1,126 @@
"""
Tests for src/font_manager.py — FontManager loading, caching, fallback,
and BDF handling, exercised against the real bundled fonts in assets/fonts.
This file replaces an earlier version whose tests were try/except blocks
ending in `assert True` — they executed the code but could not fail. Every
test here asserts observable behavior: returned font types, cache identity,
fallback selection, and BDF native-size reading.
"""
import freetype
import pytest
from unittest.mock import patch
from PIL import ImageFont
from src.font_manager import FontManager
@pytest.fixture
def mock_freetype():
"""Mock freetype module."""
with patch('src.font_manager.freetype') as mock_freetype:
yield mock_freetype
def fm():
"""A FontManager over the real assets/fonts catalog."""
return FontManager({})
class TestFontManager:
"""Test FontManager functionality."""
def test_init(self, test_config, mock_freetype):
"""Test FontManager initialization."""
# Ensure BDF files exist check passes
with patch('os.path.exists', return_value=True):
fm = FontManager(test_config)
assert fm.config == test_config
assert hasattr(fm, 'font_cache') # FontManager uses font_cache, not fonts
def test_get_font_success(self, test_config, mock_freetype):
"""Test successful font loading."""
with patch('os.path.exists', return_value=True), \
patch('os.path.join', side_effect=lambda *args: "/".join(args)):
fm = FontManager(test_config)
# Request a font (get_font requires family and size_px)
# Font may be None if font file doesn't exist in test, that's ok
try:
font = fm.get_font("small", 12) # family and size_px required
# Just verify the method can be called
assert True # FontManager.get_font() executed
except (TypeError, AttributeError):
# If method signature doesn't match, that's ok for now
assert True
def test_get_font_missing_file(self, test_config, mock_freetype):
"""Test handling of missing font file."""
with patch('os.path.exists', return_value=False):
fm = FontManager(test_config)
# Request a font where file doesn't exist
# get_font requires family and size_px
try:
font = fm.get_font("small", 12) # family and size_px required
# Font may be None if file doesn't exist, that's ok
assert True # Method executed
except (TypeError, AttributeError):
assert True # Method signature may differ
def test_get_font_invalid_name(self, test_config, mock_freetype):
"""Test requesting invalid font name."""
with patch('os.path.exists', return_value=True):
fm = FontManager(test_config)
# Request unknown font (get_font requires family and size_px)
try:
font = fm.get_font("nonexistent_font", 12) # family and size_px required
# Font may be None for unknown font, that's ok
assert True # Method executed
except (TypeError, AttributeError):
assert True # Method signature may differ
def test_get_font_with_fallback(self, test_config, mock_freetype):
"""Test font loading with fallback."""
# FontManager.get_font() requires family and size_px
# This test verifies the method exists and can be called
fm = FontManager(test_config)
assert hasattr(fm, 'get_font')
assert True # Method exists, implementation may vary
def test_load_custom_font(self, test_config, mock_freetype):
"""Test loading a custom font file directly."""
with patch('os.path.exists', return_value=True):
fm = FontManager(test_config)
# FontManager uses add_font or get_font, not load_font
# Just verify the manager can handle font operations
# The actual method depends on implementation
assert hasattr(fm, 'get_font') or hasattr(fm, 'add_font')
class TestCatalog:
def test_bundled_common_fonts_are_registered(self, fm):
# These aliases are hardcoded in FontManager.common_fonts and the
# files ship in assets/fonts — all three must resolve.
for family in ("press_start", "four_by_six", "five_by_seven"):
assert family in fm.font_catalog, f"{family} missing from catalog"
def test_catalog_families_are_lowercase_filenames(self, fm):
# _scan_fonts_directory lowercases the filename stem.
assert all(name == name.lower() for name in fm.font_catalog)
class TestGetFont:
def test_ttf_family_returns_usable_pil_font(self, fm):
font = fm.get_font("press_start", 8)
assert isinstance(font, ImageFont.FreeTypeFont)
# Usable: it can measure text.
bbox = font.getbbox("Hi")
assert bbox[2] > bbox[0]
def test_bdf_family_returns_freetype_face(self, fm):
font = fm.get_font("five_by_seven", 7)
assert isinstance(font, freetype.Face)
def test_repeat_call_returns_cached_identity(self, fm):
first = fm.get_font("press_start", 8)
hits_before = fm.performance_stats["cache_hits"]
second = fm.get_font("press_start", 8)
assert second is first
assert fm.performance_stats["cache_hits"] == hits_before + 1
def test_different_sizes_get_distinct_cache_entries(self, fm):
small = fm.get_font("press_start", 8)
large = fm.get_font("press_start", 16)
assert small is not large
assert "press_start_8" in fm.font_cache
assert "press_start_16" in fm.font_cache
def test_unknown_family_falls_back_to_default_without_raising(self, fm):
failed_before = fm.performance_stats["failed_loads"]
font = fm.get_font("no-such-family", 10)
# The documented fallback is PIL's default font (whose concrete type
# varies across Pillow versions), recorded as a failed load. It must
# still be usable for measurement.
assert type(font) is type(ImageFont.load_default())
assert font.getbbox("Hi")[2] > 0
assert fm.performance_stats["failed_loads"] == failed_before + 1
def test_corrupt_font_file_falls_back_to_default(self, fm, tmp_path):
bad = tmp_path / "broken.ttf"
bad.write_text("this is not a font file")
fm.font_catalog["broken"] = str(bad)
failed_before = fm.performance_stats["failed_loads"]
font = fm.get_font("broken", 10)
assert type(font) is type(ImageFont.load_default())
assert font.getbbox("Hi")[2] > 0
assert fm.performance_stats["failed_loads"] == failed_before + 1
class TestBdfNativeSize:
def test_five_by_seven_reports_native_height(self, fm):
# 5x7.bdf declares a 7px strike; requesting other sizes still renders
# the native size, so callers need this to know the truth.
assert fm.get_native_bdf_size("five_by_seven") == 7
def test_ttf_family_has_no_native_size(self, fm):
assert fm.get_native_bdf_size("press_start") is None
def test_unknown_family_has_no_native_size(self, fm):
assert fm.get_native_bdf_size("no-such-family") is None
class TestMeasureText:
def test_ttf_measurement_is_positive_and_cached(self, fm):
font = fm.get_font("press_start", 8)
width, height, baseline = fm.measure_text("SCORE", font)
assert width > 0 and height > 0
# Cached: same result object path on second call.
assert fm.measure_text("SCORE", font) == (width, height, baseline)
assert ("SCORE", id(font)) in fm.metrics_cache
def test_longer_text_measures_wider(self, fm):
font = fm.get_font("press_start", 8)
short, _, _ = fm.measure_text("AB", font)
long, _, _ = fm.measure_text("ABCD", font)
assert long > short
class TestCacheLifecycle:
def test_clear_cache_empties_both_caches(self, fm):
font = fm.get_font("press_start", 8)
fm.measure_text("X", font)
assert fm.font_cache and fm.metrics_cache
fm.clear_cache()
assert not fm.font_cache
assert not fm.metrics_cache
def test_reload_config_bumps_generation_and_clears(self, fm):
fm.get_font("press_start", 8)
gen_before = fm.cache_generation
fm.reload_config({})
assert fm.cache_generation == gen_before + 1
assert not fm.font_cache