refactor(cache): remove the cache layer's duplicate cleanup and dead lookups (#613)

* refactor(cache): collapse CacheStrategy's all-60 defaults table and twin soccer branch

get_sport_live_interval() without a config manager looked the sport up in
a table where every value was 60, with 60 as the fallback; it now returns
60. get_data_type_from_key() had an `if 'soccer'` branch returning the
same 'sports_live' as its else.

test_cache_strategy_intervals pins the returned strategy for every data
type x sport key x config-manager shape; it passes unchanged on the old
code. A 2,544-entry dump of every CacheStrategy method over a wider grid
is identical before and after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(cache): drop CacheStrategy's `<sport>_scoreboard` config lookup

get_sport_live_interval() and get_cache_strategy() read live/recent/
upcoming intervals from config[f"{sport}_scoreboard"]. Those sections
belonged to the built-in scoreboards the plugin system replaced; plugin
config is keyed by plugin id ("football-scoreboard"), so on a current
config the lookup always fell through to the defaults (60 live, 1800
recent, 10800 upcoming), which are now returned directly.

The one input where this differs: a config.json upgraded from the
pre-plugin era that still carries e.g. an "nfl_scoreboard" section (no
code removes them), queried with an explicit sport key. No caller in core
or the plugin monorepo passes a sport key here -- get_with_auto_strategy
only derives one for keys classed sports_live/live_scores, and its callers
(odds managers, odds-ticker) use odds keys -- so the stale section was
unreachable in practice. A dump of every CacheStrategy method over 2,544
inputs differs from the previous commit only in those 45 legacy-config
entries; the test grid now includes that shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(cache): list cache files without holding the memory-tier lock

CacheManager.list_cache_files() held the in-memory cache's lock while it
listed and stat'd the whole cache directory -- 8,864 files on a real rig
-- so every get()/set() from the display loop and plugins waited out the
scan. The lock never protected the disk: DiskCache writes and deletes
under their own lock, and a file vanishing between listdir and stat was
already handled (logged and skipped). The body is unchanged apart from
the dedent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(cache): delegate memory-tier cleanup and stats to MemoryCache

CacheManager._cleanup_memory_cache() was a line-for-line copy of
MemoryCache.cleanup(), and get_memory_cache_stats() a copy of
MemoryCache.get_stats(), both reaching into the component's private
_cache/_timestamps/_lock through "backward compatibility" aliases bound
in __init__. So the component's own cleanup and stats only ever ran in
tests, and the aliases went stale whenever the component was swapped
(test_cache_ttl_honoured does). Both now delegate, and the aliases are
gone: nothing in core, the tests, or the ledmatrix-plugins monorepo reads
them.

Behaviour is the same. Compared line by line, the two cleanups differ
only in the sort key's fallback (0 vs 0.0, which orders identically),
range+bounds check vs slice for the eviction, and the logger name on the
DEBUG summary line (src.cache_manager -> src.cache.memory_cache). A
differential run over 20,000 random memory states (str/None/garbage/
future timestamps, orphan keys, sizes 0-12, forced and throttled runs)
gives identical removed counts, resulting dicts and last-cleanup times;
the same harness catches each of three seeded mutations of
MemoryCache.cleanup. The throttle clock also moves with it:
CacheManager kept its own copy of last-cleanup, the component's is used
now, and they started equal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(background): inline the sport cache key and drop the unused request queue

get_sport_cache_key() constructed a whole CacheManager -- ConfigManager,
config parse, cache-dir probing with test-file writes -- to return
f"{sport}_{date}". It now builds the key itself in the same format as
CacheManager.generate_sport_cache_key() (UTC date, %Y%m%d); tests check
the two agree for explicit dates and, with a frozen clock at 03:30 UTC,
for the default date. Median per call on Windows: ~0.6 ms -> ~2 us
(alternating runs); on a Pi the old path also wrote a probe file per call.

request_queue was a PriorityQueue nothing ever put into: requests go
straight to the executor, so `priority` never did anything. The queue is
gone; the `priority` parameter and FetchRequest field stay (every
monorepo scoreboard passes priority=) and are documented as ignored, and
get_statistics() keeps reporting queue_size, now a literal 0 as it
always was in practice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-23 12:54:07 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent 269385c97c
commit 9d024f24ef
6 changed files with 416 additions and 188 deletions
+16 -9
View File
@@ -16,14 +16,15 @@ Key Features:
import itertools
import time
from datetime import datetime
import logging
import threading
import requests
from typing import Dict, Any, Optional, Callable, List
from dataclasses import dataclass, field
from enum import Enum
import queue
from concurrent.futures import ThreadPoolExecutor
import pytz
from src.cache_manager import CacheManager
from src.common.espn_dates import (
RANGE_RETRY_SECONDS,
@@ -57,7 +58,9 @@ class FetchRequest:
timeout: int = 30
retry_count: int = 0
max_retries: int = 3
priority: int = 1 # Higher number = higher priority
# Recorded but not acted on: requests go straight to the thread pool in
# submission order. Kept because plugins pass it through.
priority: int = 1
callback: Optional[Callable] = None
# Callbacks from submitters that JOINED this fetch instead of starting a
# duplicate one. The primary `callback` above belongs to whoever created
@@ -143,7 +146,6 @@ class BackgroundDataService:
self._request_seq = itertools.count()
self.active_requests: Dict[str, FetchRequest] = {}
self.completed_requests: Dict[str, FetchResult] = {}
self.request_queue = queue.PriorityQueue()
# Thread safety
self._lock = threading.RLock()
@@ -187,10 +189,12 @@ class BackgroundDataService:
This ensures Recent/Upcoming managers and background service
use the same cache keys.
"""
# Use the centralized cache key generation from CacheManager
from src.cache_manager import CacheManager
cache_manager = CacheManager()
return cache_manager.generate_sport_cache_key(sport, date_str)
# Same format as CacheManager.generate_sport_cache_key(). This used to
# build a whole CacheManager to call it -- config load, cache-dir
# probing with test writes -- on every submit without a cache_key.
if date_str is None:
date_str = datetime.now(pytz.utc).strftime('%Y%m%d')
return f"{sport}_{date_str}"
def submit_fetch_request(self,
sport: str,
@@ -215,7 +219,8 @@ class BackgroundDataService:
headers: HTTP headers
timeout: Request timeout
max_retries: Maximum number of retries
priority: Request priority (higher = more important)
priority: Accepted for compatibility and ignored; requests run in
submission order.
callback: Optional callback function when request completes
Returns:
@@ -719,7 +724,9 @@ class BackgroundDataService:
'completed_requests_count': len(self.completed_requests),
'max_completed_requests': self._max_completed_requests,
'completed_requests_usage_percent': (len(self.completed_requests) / self._max_completed_requests * 100) if self._max_completed_requests > 0 else 0,
'queue_size': self.request_queue.qsize(),
# Nothing is queued outside the executor; kept for callers
# that read the key.
'queue_size': 0,
'last_cleanup': self._last_completed_requests_cleanup,
'cleanup_interval': self._completed_requests_cleanup_interval
}