perf(cache): tell a stale record from its header instead of parsing it (#633)

The sports plugins cache whole season schedules: 53MB for MLB, 18MB for
NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file
takes ~1.8s with the GIL held, and every thread in the display service
waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with
the interpreter itself blocked, right on these reads. When a season record
expired, DiskCache.get paid that whole parse only to find the timestamp
too old and throw the result away.

CacheManager.set now writes timestamp and ttl ahead of the data, and
DiskCache.get reads them from the first 256 bytes of the file, applying
the same rule as before (a per-entry ttl wins over max_age; no limit
means never stale). A record that is stale is refused without being
parsed. Files in the old layout, and records from other writers, don't
match the header and are parsed in full as before.

Also: ESPN responses in the background data service and espn_dates are
parsed with orjson when it is installed (src/common/json_body.py). The
stdlib parser behind response.json() takes 3.1s on the MLB season
against orjson's 1.8s, both with the GIL held. espn_dates imports it with
a fallback, since plugins bundle copies of that module for older cores.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-24 15:50:50 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent 5baf983fe0
commit ddf5f085a5
6 changed files with 211 additions and 8 deletions
+43
View File
@@ -7,6 +7,7 @@ Handles persistent disk-based caching with atomic writes and error recovery.
import json
import math
import os
import re
import stat
import time
import tempfile
@@ -98,6 +99,40 @@ def _replace_nonfinite(obj: Any) -> Any:
# deleted. Both halves are covered by test/test_cache_nonfinite_floats.py.
#: Enough of a record to hold its header: ``{"timestamp":<float>,"ttl":<n>,``.
_HEAD_BYTES = 256
#: A record written with its header first (CacheManager.set does). Anything
#: else -- older files with "data" first, records from other writers -- does not
#: match and is parsed in full, as before.
_HEAD_RE = re.compile(
rb'\A\s*\{\s*"timestamp"\s*:\s*(-?[0-9][0-9.eE+-]*)\s*'
rb'(?:,\s*"ttl"\s*:\s*(-?[0-9][0-9.eE+-]*))?\s*[,}]'
)
def _stale_from_head(head: bytes, max_age: Optional[int], now: float) -> bool:
"""True when a record's header alone shows it has expired.
Mirrors the expiry rule in DiskCache.get: a per-entry ttl wins over the
caller's max_age, and no limit at all means never stale. False whenever the
header cannot be read, so the full parse decides as it always did.
"""
match = _HEAD_RE.match(head)
if not match:
return False
try:
timestamp = float(match.group(1))
limit = max_age
if match.group(2) is not None:
ttl = float(match.group(2))
if ttl >= 0:
limit = ttl
except ValueError:
return False
return limit is not None and (now - timestamp) > limit
if orjson is not None:
# Encoding the cache record dominated the background fetch worker: on a
# Pi 4, stdlib json.dumps runs ~12ms per MB and holds the GIL for all of
@@ -266,6 +301,14 @@ class DiskCache:
try:
with self._lock:
with open(cache_path, 'rb') as f:
# Decide staleness from the header before paying for the
# parse. A stale read is the common case for the biggest
# records (a season schedule is re-fetched when its cache
# expires), and parsing 53MB to throw it away held the GIL
# for ~1.8s -- a visible freeze on the panel.
if _stale_from_head(f.read(_HEAD_BYTES), max_age, time.time()):
return None
f.seek(0)
record = _loads(f.read())
# Determine record timestamp (prefer embedded, else file mtime)