refactor(web): one thread-safe TTL cache for the web process

web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.

Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
  counted and get_cached defaulted to 60s. An entry now expires after the TTL
  it was stored with; a reader's ttl_seconds can only shorten that. Both
  current callers pass the same value on both sides (fonts_catalog 300s,
  system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
  same expired key could raise KeyError (reproduced), which the endpoints
  turned into a 500.

app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.

Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-23 16:59:06 -04:00
co-authored by Claude Opus 5.5
parent 70eab6fb25
commit 7789275499
3 changed files with 263 additions and 63 deletions
+28 -31
View File
@@ -295,34 +295,41 @@ try:
except ImportError:
pass
# Cached AP mode check — avoids creating a WiFiManager per request
_ap_mode_cache = {'value': False, 'timestamp': 0}
# systemctl answers, memoised so they are not a subprocess fork per request
# (AP mode) or per SSE tick (display service). A failed check keeps the last
# known answer for the same TTL rather than retrying on every request.
from web_interface.cache import TTLCache
_service_status_cache = TTLCache()
_AP_MODE_CACHE_TTL = 30 # seconds — AP mode is user-initiated; 30s is fine
# Cached ledmatrix service status for SSE stats stream
_ledmatrix_service_cache = {'active': False, 'timestamp': 0}
_LEDMATRIX_SERVICE_CACHE_TTL = 15 # seconds
def _unit_is_active(unit, ttl):
"""`systemctl is-active <unit>`, cached for ``ttl`` seconds.
False where there is no systemctl (a dev machine); on a failed check, the
last known answer.
"""
active = _service_status_cache.get(unit)
if active is not None:
return active
active = _service_status_cache.peek(unit, False)
if _SYSTEMCTL:
try:
result = subprocess.run([_SYSTEMCTL, 'is-active', unit],
capture_output=True, text=True, timeout=2)
active = result.stdout.strip() == 'active'
except (subprocess.SubprocessError, OSError) as e:
logging.getLogger('web_interface').warning(
"systemctl is-active %s failed: %s", unit, e)
_service_status_cache.set(unit, active, ttl=ttl)
return active
def is_ap_mode_active():
"""
Check if access point mode is currently active (cached, 30s TTL).
Uses a direct systemctl check instead of instantiating WiFiManager.
"""
now = time.time()
if (now - _ap_mode_cache['timestamp']) < _AP_MODE_CACHE_TTL:
return _ap_mode_cache['value']
try:
result = subprocess.run(
['systemctl', 'is-active', 'hostapd'],
capture_output=True, text=True, timeout=2
)
active = result.stdout.strip() == 'active'
_ap_mode_cache['value'] = active
_ap_mode_cache['timestamp'] = now
return active
except (subprocess.SubprocessError, OSError) as e:
logging.getLogger('web_interface').error(f"AP mode check failed: {e}")
return _ap_mode_cache['value']
return _unit_is_active('hostapd', _AP_MODE_CACHE_TTL)
# Captive portal detection endpoints
# When AP mode is active, return responses that TRIGGER the captive portal popup.
@@ -672,17 +679,7 @@ def system_status_generator():
cpu_temp = metrics['cpu_temp']
# Check if display service is running (cached to avoid per-client subprocess forks)
now = time.time()
if (now - _ledmatrix_service_cache['timestamp']) >= _LEDMATRIX_SERVICE_CACHE_TTL:
if _SYSTEMCTL:
try:
result = subprocess.run([_SYSTEMCTL, 'is-active', 'ledmatrix'],
capture_output=True, text=True, timeout=2)
_ledmatrix_service_cache['active'] = result.stdout.strip() == 'active'
except (subprocess.SubprocessError, OSError) as e:
app.logger.warning("systemctl status check failed: %s", e)
_ledmatrix_service_cache['timestamp'] = now
service_active = _ledmatrix_service_cache['active']
service_active = _unit_is_active('ledmatrix', _LEDMATRIX_SERVICE_CACHE_TTL)
status = {
'timestamp': time.time(),