mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
* fix(errors): record the exception's own stack trace record_error() called traceback.format_exc(), which only sees an exception while its except block is running. plugin_executor records exceptions caught on a worker thread after that block has ended, so every trace on /errors read "NoneType: None". The trace is now built from the exception's __traceback__. The executor's log call had the same problem with exc_info=True and now passes the exception. record_error() also merged LEDMatrixError context into the caller's dict in place; it now works on a copy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(wifi): point at configure_wifi_permissions.sh instead of a sudoers list The module docstring told users to grant NOPASSWD sudo on iptables and ip. configure_wifi_permissions.sh refuses those grants on purpose: a wildcard rule for either runs an arbitrary program as root. Point at the script and say why it leaves them out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): disconnect finds the saved profile by SSID disconnect_from_network() asked `nmcli -f NAME,802-11-wireless.ssid connection show` for the profile to take down, but nmcli rejects that column for `connection show`, so the lookup always failed and only the device was disconnected. The per-profile lookup _connect_nmcli() already used is now _find_profile_for_ssid(), and both callers share it. It also splits terse output on the last colon and unescapes "\:", so a profile name containing a colon is found. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): write wifi_config.json atomically and report a failed save _save_config() opened the file for writing in place and swallowed any error, so a wifi_config.json left owned by root made the web toggle for auto-enabling AP mode report success while nothing was saved, and a crash mid-write could truncate the file. It now uses atomic_write_json, which also keeps the file's owner and shared group when root saves it, and returns False on failure. POST /wifi/ap/auto-enable answers 500 in that case. The file is now written with indent=4, like the other config files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): resolve plugin:// fonts in the plugin's own directory FontManager looked for a plugin's bundled fonts under Path("plugins") / plugin_id: relative to the process cwd, and not the default install directory (plugin-repos/), so a manifest's plugin:// fonts never loaded. register_plugin_fonts() takes an optional plugin_dir, and PluginManager passes the directory it loaded the plugin from. Callers that omit it get a lookup in the configured plugin_system.plugins_directory, then plugins/, resolved against the install root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(api-helper): cache responses for the requested cache_ttl APIHelper.get(cache_ttl=...) and set_cache(ttl=...) dropped the ttl on the claim that CacheManager does not support one, but CacheManager.set() takes a ttl, stores it with the entry, and both cache tiers honour it over a reader's max_age. Without it every response expired after the 300-second default read age, whatever the plugin asked for. The ttl is now passed through, and the cache read passes cache_ttl as max_age for entries written without one. The class docstring describes what the helper actually does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(style): one scale range for the schema, element_scale and LogoHelper The generated Scale field allowed 0.1 to 10, element_style's reader capped at 10 with no floor, and LogoHelper accepted 0.05 to 8 and reset anything else to 1.0. A logo scale of 9, which the form accepts, drew at the shipped size. MIN_ELEMENT_SCALE / MAX_ELEMENT_SCALE (0.1, 10.0) in src.element_style are now the schema bounds and the clamp every reader applies through coerce_scale(): a positive number outside the range is clamped, and anything that is not a finite positive number means the default. That also stops element_scale() passing NaN through, since min(nan, 10.0) is nan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(logos): placeholder lands at the requested path; empty logos list download_missing_logo() wrote its fallback placeholder to <normalize_abbreviation(abbr)>.png in the logo directory rather than to the logo_path the caller passed, so it could return True while nothing existed where the plugin looks (e.g. "TA&M.png" vs "TAANDM.png"). create_placeholder_logo() takes an optional filepath, and download_missing_logo passes the requested one. download_missing_logo_for_team() only caught KeyError, so a team whose "logos" list is empty raised IndexError; it now treats KeyError, IndexError and TypeError as "no logo URL". The placeholder is drawn with PLACEHOLDER_SIZE / PLACEHOLDER_BG, the constants is_placeholder_logo() recognises it by, instead of repeated literals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): resolve bundled font paths against the install root TextHelper's default font_dir, the logo placeholder's font and FontManager's font_overrides.json were all relative to the process cwd, so a process started anywhere but the install root (the plugin safety harness, a manual run, a unit without WorkingDirectory) drew with PIL's default face and read no overrides. They now go through font_layout.resolve_asset_path; the overrides file sits in the install root's config/. The resolver docstrings described an order the code does not follow: resolve_asset_path never consults the cwd, and sports_shared's _resolve_font_path tries the cwd first. Both docstrings now say what the code does, and _resolve_font_path calls resolve_asset_path instead of probing FontManager for it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(sync): the web UI reads the sync status file the display writes sync_manager writes its status to tempfile.gettempdir(), but GET /api/v3/sync/status read a hardcoded /tmp/led_matrix_sync_status.json and defaulted the port to a literal 5765. Wherever TMPDIR is set (or on any non-/tmp host) the page only ever showed "starting". The endpoint now uses sync_manager.STATUS_FILE and SYNC_PORT. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(http): the rankings resolver sends the project's User-Agent DynamicTeamResolver fetched ESPN rankings with a bare requests.get, so it sent python-requests' default User-Agent, which ESPN rejects; the AP_TOP_N favourites then resolved to nothing. It now sends DEFAULT_HTTP_HEADERS. BaseOddsManager carried its own copy of the User-Agent string and now uses the same shared headers (which also adds Accept-Language). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(backup): record the core release and read the configured plugin dir The manifest's ledmatrix_version came from a VERSION file that does not exist, then from .git/HEAD: a 12-character sha, or "ref: refs/he" when the branch's ref was packed. It is now src.__version__. list_installed_plugins() scanned a hardcoded plugin-repos/, so on an install whose plugin_system.plugins_directory points elsewhere, plugins missing from plugin_state.json were left out of the backup. It now reads the configured directory from config/config.json, defaulting to plugin-repos. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(startup): report a missing display section once A config without a display section produced three errors for the one problem ("Missing required configuration key: display", "Display configuration is missing or empty" and "Display configuration is missing"), and an empty one produced two. _validate_config now reports it once, as a missing key or an empty section, and _validate_display_config leaves it to that. The module docstring said the validator fails fast; nothing in the display service calls raise_on_errors(), so it now says the errors are reported and startup continues. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(wifi): share the copied blocks and name the AP constants - _parse_nmcli_wifi_list() is the one parser behind _scan_nmcli and _scan_nmcli_cached. - _verify_connected(), _wait_for_device_idle(), _failsafe_ap() and _mark_forced() replace blocks that were pasted two or three times in the connect and enable-AP paths. The device-idle wait now checks before its first one-second sleep instead of after it. - _check_command() calls _find_command_path() instead of repeating it. - AP_IP, PORTAL_PORT, AP_PROFILE_NAME and AP_PROFILE_NAMES name values that were spelled out 14, 12, 8 and 2 times; the two deletion loops now walk the same tuple. The iwconfig status path compares the AP address exactly: startswith() also skipped 192.168.4.10-19. - Dropped a second WIFI.SIGNAL query that repeated the first, a no-op "if ssid: continue", the try/except around _connect_wpa_supplicant's constant return, and a second save of a scan scan_networks already saves. - _ensure_wifi_radio_enabled's docstring says it returns True when the radio state cannot be read at all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(config): drop dead branches and history comments in ConfigManager - The module docstring pointed plugin authors at update_plugin_config(), which does not exist; it now names save_config_atomic() and save_raw_file_content(). - load_config's FileNotFoundError handler tested the message for "config_secrets.json", but a missing secrets file is handled where it is read, so only config.json reaches it; the check is gone. - save_raw_file_content's `file_type == "main" or "secrets"` guard was always true (anything else raised earlier). - get_raw_file_content('secrets') already returns {} for a missing file, so the os.path.exists() in front of two calls to it is gone. - Comments that narrated earlier behaviour are rewritten as what the code does now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(background-data): present-tense comments, drop unused API - Comments that told the history of each fix (what "used to" happen, "the old per-delivery release") now state the invariant the code keeps. - get_statistics() no longer reports a constant 'queue_size': 0, and the uncalled clear_completed_requests() is gone (_cleanup_completed_requests does that job on every completion). Neither is referenced in core, the web UI or the plugin monorepo. shutdown_background_service() has no production caller either, but it is the only way to tear down the get_background_service() singleton, which the tests rely on, so it stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(odds): drop the unread cache_ttl and merge the odds_data branches BaseOddsManager loaded base_odds_manager.cache_ttl from config and never used it: cached odds live for the update interval (get_odds' ttl=interval). No core or monorepo code reads the attribute, so it is gone along with its log line. The two consecutive `if odds_data:` blocks are one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(backup): one table for the single-file sections config, secrets, wifi and ytm_auth were each spelled out in create, preview, validate and restore. _SINGLE_FILE_SECTIONS lists them once, with the RestoreOptions flag that restores each, and all four walk it. Restore error messages keep their wording ("Failed to restore <file name>"). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(fonts): drop FontManager's write-only state and duplicate logs - fonts_config, font_metadata and font_dependencies were written and never read; the performance_stats keys font_load_times, render_times, total_renders and the per-call "resolve" timings (_record_performance_metric) likewise. get_performance_stats() reads only the counters that remain. Nothing in core or the plugin monorepo references any of them. - A failed BDF load was logged twice, by _load_bdf_font and again by get_font; get_font's line is the one kept. - Removed "NEW:" and commented-out cozette entries, the "Copy font to assets/fonts" comment on code that copies nothing, and local imports of names the module already imports. The deprecated add_font() now resolves assets/fonts against the install root. The @deprecated methods stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(text-helper): cache loaded fonts; drop the pre-textlength fallback TextHelper declared _font_cache, cleared it and reported its size, but never stored anything in it. load_fonts() now keeps each (file, size) it loads there, so clear_font_cache() and get_font_cache_stats() mean what they say and repeated load_fonts() calls reuse the fonts. get_text_width() no longer catches AttributeError for Pillow releases without ImageDraw.textlength; requirements.txt pins Pillow>=12.2. The class docstring describes what the helper does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(common): fix wrong docstrings in api_helper, permission_utils, snapshot_policy - permission_utils called 0o2775 "sticky bit"; the 2 is setgid, which is what makes new files take the directory's group. - snapshot_policy pointed at web_interface/blueprints/api_v3.py, which is a package now; the health check is in api_v3/misc.py. - APIHelper.clear_cache() lost a history note and a fallback to a clear() method that neither CacheManager nor the testing MockCacheManager has. The session headers are built from DEFAULT_HTTP_HEADERS instead of a copy of them, and the module docstring says what the module offers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(sports): present-tense comments in the shared scoreboard renderers - sports_scroll and sports_game_renderer comments that referred to "this PR", "the old flat 128px card" or what the renderer "previously" did now describe the current behaviour and its reason. - The block explaining why non-finite settings are rejected sat above _score_reserve_width; it describes _center_gap_width and now lives in it. - unshare_element_fonts wrapped its import of font_layout.load_truetype in an `except ImportError` that cannot fire inside core; the import stays at call time so tests can spy on the pinned loader. - sports_card docstrings that told the history of a fix say what the code does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sports-shared): drop dead code, name the ESPN limit - _get_weeks_data asked for limit=1000, which fetch_espn_scoreboard clamps to ESPN_MAX_LIMIT anyway; it now names that constant. Its unused `immediate_events = []` is gone. - _get_season_schedule_dates() returned ("", "") and has no caller in core or the plugin monorepo. - _should_log keeps its warning_type parameter (part of the inherited signature, though nothing in core or the monorepo calls it) and its docstring says the cooldown is shared across types. - An unused ImageFont import is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sync): one follower-mode switch, shared panel defaults - The class docstring said the leader sends PNG frames. Frames go over UDP as raw RGB; PNG is only the Vegas scroll image sent over TCP. It now describes both paths. - _enter_follower_mode() replaces the two copies of "note the leader, switch from standalone to follower, log, write status" in the frame and scroll-position handlers. - The rows/cols fallbacks use DEFAULT_ROWS / DEFAULT_COLS from src.display_geometry, as chain_length already did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(style): drop _layout_axis, name the layout group title - ElementStyleResolver._layout_axis() had no caller in core or the plugin monorepo. - _element_block_from_spec checked spec['size'] was a dict again after size_spec already had; it reads size_spec. - The "Layout Offsets" title written into three generated schema blocks is _LAYOUT_TITLE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(logo-helper): say what the placeholder draws; name the 1.5 box factor - _create_placeholder_logo's docstring said it draws the team abbreviation; it draws an outlined grey box and nothing else. The docstring says so, and the "in a real implementation you'd want text" comments are gone. - The 1.5 x panel default logo box, written out six times, is DEFAULT_LOGO_BOX_FACTOR. - ImageDraw is imported with Image at the top of the module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(logos): drop dead code and a duplicate regex in logo_downloader - _SAFE_LEAGUE_CODE_RE was the same pattern as _SAFE_LEAGUE_RE; both checks use the one. - get_logo_filename_variations reassigned the TA&M case to the list it already had; the function returns the two names directly. - _get_team_name_variations() had no caller in core or the plugin monorepo. - fetch_single_team's docstring was copied from fetch_teams_data; a log message read "for{team_id}". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: drop the Pillow<9.1 resample shim and a catch-and-reraise - adaptive_images fell back to Image.LANCZOS/NEAREST for Pillow < 9.1; requirements.txt pins Pillow>=12.2. RESAMPLE_LANCZOS and RESAMPLE_NEAREST keep their names (src.common re-exports them). - CacheManager.save_cache caught CacheError only to re-raise it; the disk write is now called directly, with the same result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(api-helper): stop the real CacheManager's cleanup thread The cache-lifetime tests built a CacheManager and left its cleanup thread's class-wide claim on the directory in place, which broke test_cache_cleanup_thread_ownership when it ran later in the session. The fixture now stops the thread on teardown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): core-common Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
836 lines
34 KiB
Python
836 lines
34 KiB
Python
"""
|
|
Background Data Service for LEDMatrix
|
|
|
|
This service provides background threading capabilities for season data fetching
|
|
to prevent blocking the main display loop. It's designed to be used across
|
|
all sport managers for consistent background data management.
|
|
|
|
Key Features:
|
|
- Thread-safe data caching
|
|
- Automatic retry logic with exponential backoff
|
|
- Configurable timeouts and intervals
|
|
- Graceful error handling
|
|
- Progress tracking and logging
|
|
- Memory-efficient data storage
|
|
"""
|
|
|
|
import itertools
|
|
import time
|
|
from datetime import datetime
|
|
import logging
|
|
import threading
|
|
import requests
|
|
from typing import Dict, Any, Optional, Callable, List
|
|
from dataclasses import dataclass, field
|
|
from enum import Enum
|
|
from concurrent.futures import ThreadPoolExecutor
|
|
import pytz
|
|
from src.cache_manager import CacheManager
|
|
from src.common.json_body import response_json
|
|
from src.common.espn_dates import (
|
|
RANGE_RETRY_SECONDS,
|
|
_note_range_rejected,
|
|
_ranges_known_rejected,
|
|
clamp_espn_limit,
|
|
fetch_espn_date_chunks,
|
|
parse_espn_date_range,
|
|
)
|
|
# Configure logging
|
|
logger = logging.getLogger(__name__)
|
|
|
|
class FetchStatus(Enum):
|
|
"""Status of background fetch operations."""
|
|
PENDING = "pending"
|
|
IN_PROGRESS = "in_progress"
|
|
COMPLETED = "completed"
|
|
FAILED = "failed"
|
|
CANCELLED = "cancelled"
|
|
|
|
@dataclass
|
|
class FetchRequest:
|
|
"""Represents a background fetch request."""
|
|
id: str
|
|
sport: str
|
|
year: int
|
|
cache_key: str
|
|
url: str
|
|
params: Dict[str, Any] = field(default_factory=dict)
|
|
headers: Dict[str, str] = field(default_factory=dict)
|
|
timeout: int = 30
|
|
retry_count: int = 0
|
|
max_retries: int = 3
|
|
# Recorded but not acted on: requests go straight to the thread pool in
|
|
# submission order. Kept because plugins pass it through.
|
|
priority: int = 1
|
|
callback: Optional[Callable] = None
|
|
# Callbacks from submitters that JOINED this fetch instead of starting a
|
|
# duplicate one. The primary `callback` above belongs to whoever created
|
|
# the request; these belong to everyone who asked for the same cache_key
|
|
# while it was still in flight.
|
|
extra_callbacks: List[Callable] = field(default_factory=list)
|
|
created_at: float = field(default_factory=time.time)
|
|
status: FetchStatus = FetchStatus.PENDING
|
|
# Set once the worker has decided this response will be cached, while it
|
|
# still holds the lock. From that point cancelling is refused: the write
|
|
# is already authorised, and abandoning it here would put the payload in
|
|
# the cache with the callbacks suppressed -- joiners waiting forever for a
|
|
# fetch that did, in fact, succeed.
|
|
commit_claimed: bool = False
|
|
result: Optional[Any] = None
|
|
error: Optional[str] = None
|
|
|
|
@dataclass
|
|
class FetchResult:
|
|
"""Result of a background fetch operation.
|
|
|
|
``data`` survives on the stored result only for requests submitted without
|
|
a ``callback``, where polling ``get_result()`` is the sole way to collect
|
|
it. When a callback was given, the payload has already been delivered and
|
|
the service releases it -- see :meth:`BackgroundDataService._release_payload`.
|
|
Either way the data remains in the cache under the request's ``cache_key``,
|
|
which is where consumers read it from.
|
|
"""
|
|
request_id: str
|
|
success: bool
|
|
data: Optional[Any] = None
|
|
error: Optional[str] = None
|
|
cached: bool = False
|
|
fetch_time: float = 0.0
|
|
retry_count: int = 0
|
|
completed_at: float = field(default_factory=time.time) # Timestamp when request completed
|
|
# The request's final status, recorded so a finished request can still be
|
|
# reported accurately. Without it a caller can only be told COMPLETED or
|
|
# FAILED, which turns "you cancelled this" into "this errored".
|
|
final_status: Optional[FetchStatus] = None
|
|
|
|
class BackgroundDataService:
|
|
"""
|
|
Background data service for fetching season data without blocking the main thread.
|
|
|
|
This service manages a pool of background threads to fetch data asynchronously,
|
|
with intelligent caching, retry logic, and progress tracking.
|
|
"""
|
|
|
|
# Plugins feature-detect this. A core without it sends season ranges to
|
|
# ESPN as-is and gets 400s since 2026-09-15, so plugins fetch those
|
|
# ranges themselves instead of submitting them here.
|
|
handles_espn_date_ranges = True
|
|
|
|
def __init__(self, cache_manager: CacheManager, max_workers: int = 3, request_timeout: int = 30):
|
|
"""
|
|
Initialize the background data service.
|
|
|
|
Args:
|
|
cache_manager: Cache manager instance for storing fetched data
|
|
max_workers: Maximum number of background threads
|
|
request_timeout: Default timeout for HTTP requests
|
|
"""
|
|
self.cache_manager = cache_manager
|
|
self.max_workers = max_workers
|
|
self.request_timeout = request_timeout
|
|
|
|
# Thread management
|
|
self.executor = ThreadPoolExecutor(max_workers=max_workers, thread_name_prefix="BackgroundData")
|
|
# cache_key -> request_id for fetches currently in flight, so a second
|
|
# submit for the same key joins the running fetch instead of starting
|
|
# another. It is the normal case: a sport's Recent and Upcoming
|
|
# managers miss the cache for the same season schedule together.
|
|
self._inflight_by_cache_key: Dict[str, str] = {}
|
|
# Makes every request_id unique. The id also carries a millisecond
|
|
# timestamp, but two submits can share a millisecond, and a joiner
|
|
# uses the id as its handle for get_result().
|
|
self._request_seq = itertools.count()
|
|
self.active_requests: Dict[str, FetchRequest] = {}
|
|
self.completed_requests: Dict[str, FetchResult] = {}
|
|
|
|
# Thread safety
|
|
self._lock = threading.RLock()
|
|
self._shutdown = False
|
|
|
|
# Cleanup tracking
|
|
self._max_completed_requests = 500 # Maximum completed requests to keep
|
|
self._completed_requests_cleanup_interval = 600.0 # Cleanup every 10 minutes
|
|
self._last_completed_requests_cleanup = time.time()
|
|
|
|
# Statistics
|
|
self.stats = {
|
|
'total_requests': 0,
|
|
'completed_requests': 0,
|
|
'failed_requests': 0,
|
|
'cached_hits': 0,
|
|
'cache_misses': 0,
|
|
'total_fetch_time': 0.0,
|
|
'average_fetch_time': 0.0
|
|
}
|
|
|
|
# Session for HTTP requests
|
|
self.session = requests.Session()
|
|
self.session.mount('http://', requests.adapters.HTTPAdapter(max_retries=3))
|
|
self.session.mount('https://', requests.adapters.HTTPAdapter(max_retries=3))
|
|
|
|
# Default headers: core's shared set (real User-Agent, no hand-set
|
|
# Accept-Encoding) -- see src/common/api_helper.py.
|
|
from src.common.api_helper import DEFAULT_HTTP_HEADERS
|
|
self.default_headers = dict(DEFAULT_HTTP_HEADERS)
|
|
|
|
logger.info(f"BackgroundDataService initialized with {max_workers} workers")
|
|
|
|
def get_sport_cache_key(self, sport: str, date_str: str = None) -> str:
|
|
"""
|
|
Generate consistent cache keys for sports data.
|
|
This ensures Recent/Upcoming managers and background service
|
|
use the same cache keys.
|
|
"""
|
|
# Same format as CacheManager.generate_sport_cache_key(), built here
|
|
# rather than by constructing a CacheManager (config load, cache-dir
|
|
# probing) on every submit without a cache_key.
|
|
if date_str is None:
|
|
date_str = datetime.now(pytz.utc).strftime('%Y%m%d')
|
|
return f"{sport}_{date_str}"
|
|
|
|
def submit_fetch_request(self,
|
|
sport: str,
|
|
year: int,
|
|
url: str,
|
|
cache_key: str = None,
|
|
params: Optional[Dict[str, Any]] = None,
|
|
headers: Optional[Dict[str, str]] = None,
|
|
timeout: Optional[int] = None,
|
|
max_retries: int = 3,
|
|
priority: int = 1,
|
|
callback: Optional[Callable] = None) -> str:
|
|
"""
|
|
Submit a background fetch request.
|
|
|
|
Args:
|
|
sport: Sport identifier (e.g., 'nfl', 'ncaafb')
|
|
year: Year to fetch data for
|
|
url: URL to fetch data from
|
|
cache_key: Cache key for storing/retrieving data
|
|
params: URL parameters
|
|
headers: HTTP headers
|
|
timeout: Request timeout
|
|
max_retries: Maximum number of retries
|
|
priority: Accepted for compatibility and ignored; requests run in
|
|
submission order.
|
|
callback: Optional callback function when request completes
|
|
|
|
Returns:
|
|
Request ID for tracking the fetch operation
|
|
"""
|
|
if self._shutdown:
|
|
raise RuntimeError("BackgroundDataService is shutting down")
|
|
|
|
# Generate cache key if not provided
|
|
if cache_key is None:
|
|
cache_key = self.get_sport_cache_key(sport)
|
|
|
|
with self._lock:
|
|
request_id = (f"{sport}_{year}_{int(time.time() * 1000)}"
|
|
f"_{next(self._request_seq)}")
|
|
|
|
# Check cache first
|
|
cached_data = self.cache_manager.get(cache_key)
|
|
if cached_data:
|
|
with self._lock:
|
|
self.stats['cached_hits'] += 1
|
|
result = FetchResult(
|
|
request_id=request_id,
|
|
success=True,
|
|
data=cached_data,
|
|
cached=True,
|
|
fetch_time=0.0
|
|
)
|
|
# Filed before the callback runs, as it always was: a callback
|
|
# that queries get_result()/is_request_complete() for its own
|
|
# request must still find it. Releasing afterwards mutates the
|
|
# same object the dict holds.
|
|
self.completed_requests[request_id] = result
|
|
|
|
if callback:
|
|
try:
|
|
callback(result)
|
|
except Exception as e:
|
|
logger.error(f"Error in callback for request {request_id}: {e}")
|
|
self._release_payload(result)
|
|
|
|
logger.debug(f"Cache hit for {sport} {year} data")
|
|
return request_id
|
|
|
|
# limit above 500 makes an ESPN *scoreboard* return a truncated list
|
|
# (src/common/espn_dates.py). Other endpoints need more: /teams has 762
|
|
# college-football teams, so only scoreboards are clamped.
|
|
if url.split('?', 1)[0].rstrip('/').endswith('/scoreboard'):
|
|
params = clamp_espn_limit(params)
|
|
|
|
# Create fetch request
|
|
request = FetchRequest(
|
|
id=request_id,
|
|
sport=sport,
|
|
year=year,
|
|
cache_key=cache_key,
|
|
url=url,
|
|
params=dict(params or {}),
|
|
headers={**self.default_headers, **(headers or {})},
|
|
timeout=timeout or self.request_timeout,
|
|
max_retries=max_retries,
|
|
priority=priority,
|
|
callback=callback
|
|
)
|
|
|
|
with self._lock:
|
|
existing_id = self._inflight_by_cache_key.get(cache_key)
|
|
existing = self.active_requests.get(existing_id) if existing_id else None
|
|
if existing_id and existing is None:
|
|
# Stranded index entry: the request it names is gone. Drop it and
|
|
# fetch normally. Looking the request up rather than trusting the
|
|
# id is what stops a stale entry wedging a key forever.
|
|
del self._inflight_by_cache_key[cache_key]
|
|
if existing is not None:
|
|
# Someone is already fetching this key. Ride along rather than
|
|
# duplicating the download, the parse and the resident copy.
|
|
if callback:
|
|
existing.extra_callbacks.append(callback)
|
|
self.stats['deduplicated_requests'] = (
|
|
self.stats.get('deduplicated_requests', 0) + 1
|
|
)
|
|
logger.info(
|
|
"Joined in-flight fetch %s for %s (cache_key=%s) instead of "
|
|
"starting a duplicate", existing_id, sport, cache_key
|
|
)
|
|
return existing_id
|
|
|
|
self.active_requests[request_id] = request
|
|
self._inflight_by_cache_key[cache_key] = request_id
|
|
self.stats['total_requests'] += 1
|
|
self.stats['cache_misses'] += 1
|
|
|
|
# Submit to executor
|
|
self.executor.submit(self._fetch_data_worker, request)
|
|
|
|
logger.info(f"Submitted background fetch request {request_id} for {sport} {year}")
|
|
return request_id
|
|
|
|
def _fetch_data_worker(self, request: FetchRequest) -> FetchResult:
|
|
"""
|
|
Worker function that performs the actual data fetching.
|
|
|
|
Args:
|
|
request: Fetch request to process
|
|
|
|
Returns:
|
|
Fetch result with data or error information
|
|
"""
|
|
start_time = time.time()
|
|
result = FetchResult(request_id=request.id, success=False, retry_count=request.retry_count)
|
|
|
|
try:
|
|
with self._lock:
|
|
# A request cancelled while it sat in the executor queue stays
|
|
# cancelled: no download, no cache write, no callback.
|
|
if request.status == FetchStatus.CANCELLED:
|
|
cancelled_before_start = True
|
|
else:
|
|
cancelled_before_start = False
|
|
request.status = FetchStatus.IN_PROGRESS
|
|
if cancelled_before_start:
|
|
logger.info(
|
|
"Request %s was cancelled before its worker started; "
|
|
"skipping the fetch entirely", request.id
|
|
)
|
|
# Assign before returning: the finally block stores `result`
|
|
# in completed_requests, so building a fresh one here would
|
|
# file the untouched placeholder instead of this outcome.
|
|
result = FetchResult(
|
|
request_id=request.id,
|
|
success=False,
|
|
error="cancelled",
|
|
fetch_time=time.time() - start_time,
|
|
retry_count=request.retry_count
|
|
)
|
|
return result
|
|
|
|
logger.info(f"Starting background fetch for {request.sport} {request.year}")
|
|
|
|
# ESPN stopped accepting dates=YYYYMMDD-YYYYMMDD on 2026-09-15 and
|
|
# answers 400 for every sport. Re-ask in months and days rather
|
|
# than let a whole season fail. See src/common/espn_dates.py.
|
|
# The "ranges are rejected" memo is shared with
|
|
# fetch_espn_scoreboard(): once either path has seen a range
|
|
# rejected, the other skips the doomed range request too.
|
|
is_range = parse_espn_date_range(request.params.get("dates")) is not None
|
|
data = None
|
|
chunks_tried = False
|
|
if is_range and _ranges_known_rejected():
|
|
data = self._fetch_in_date_chunks(request)
|
|
# Every chunk failed: ask for the range itself below so the
|
|
# failure carries a real HTTP error, without re-spending chunks.
|
|
chunks_tried = data is None
|
|
|
|
if data is None:
|
|
# Perform HTTP request with retry logic
|
|
response = self._make_request_with_retry(request)
|
|
if is_range and response.status_code == 400 and not chunks_tried:
|
|
_note_range_rejected()
|
|
logger.warning(
|
|
"ESPN rejected the date range %s (400); fetching it as "
|
|
"month/day chunks, and fetching ranges that way for the "
|
|
"next %d hours",
|
|
request.params.get("dates"), RANGE_RETRY_SECONDS // 3600,
|
|
)
|
|
data = self._fetch_in_date_chunks(request)
|
|
if data is None:
|
|
response.raise_for_status()
|
|
else:
|
|
response.raise_for_status()
|
|
data = response_json(response)
|
|
|
|
# Validate data structure
|
|
if not isinstance(data, dict):
|
|
raise ValueError(f"Expected dict response, got {type(data)}")
|
|
|
|
if 'events' not in data:
|
|
raise ValueError("Response missing 'events' field")
|
|
|
|
# Validate events structure
|
|
events = data.get('events', [])
|
|
if not isinstance(events, list):
|
|
raise ValueError(f"Expected events to be list, got {type(events)}")
|
|
|
|
# Log data validation
|
|
logger.debug(f"Validated {len(events)} events for {request.sport} {request.year}")
|
|
|
|
# A cancelled request must not commit anything. Cancelling
|
|
# releases the cache_key, so a replacement fetch for the same key
|
|
# may already be in flight or finished -- writing this response to
|
|
# the cache now would overwrite fresher data with the response
|
|
# nobody wanted. The worker has no way to abort the HTTP call, so
|
|
# this is where the work gets discarded.
|
|
with self._lock:
|
|
cancelled = request.status == FetchStatus.CANCELLED
|
|
if not cancelled:
|
|
# Claim the commit in the same critical section that read
|
|
# the status, so a cancel cannot slip in between the check
|
|
# and the cache write below. The write itself stays outside
|
|
# the lock: it serialises a multi-megabyte payload to the
|
|
# SD card, and holding the service lock across that would
|
|
# stall every submit, status query and cancel behind it.
|
|
request.commit_claimed = True
|
|
if cancelled:
|
|
logger.info(
|
|
"Discarding response for cancelled request %s; %s may "
|
|
"already belong to a replacement fetch",
|
|
request.id, request.cache_key
|
|
)
|
|
result = FetchResult(
|
|
request_id=request.id,
|
|
success=False,
|
|
error="cancelled",
|
|
fetch_time=time.time() - start_time,
|
|
retry_count=request.retry_count
|
|
)
|
|
return result
|
|
|
|
# Cache the data
|
|
self.cache_manager.set(request.cache_key, data)
|
|
|
|
# Update request status
|
|
with self._lock:
|
|
request.status = FetchStatus.COMPLETED
|
|
request.result = data
|
|
|
|
# Create successful result
|
|
fetch_time = time.time() - start_time
|
|
result = FetchResult(
|
|
request_id=request.id,
|
|
success=True,
|
|
data=data,
|
|
fetch_time=fetch_time,
|
|
retry_count=request.retry_count
|
|
)
|
|
|
|
logger.info(f"Successfully fetched {request.sport} {request.year} data in {fetch_time:.2f}s")
|
|
|
|
except Exception as e:
|
|
error_msg = str(e)
|
|
logger.error(f"Failed to fetch {request.sport} {request.year} data: {error_msg}")
|
|
|
|
with self._lock:
|
|
# A cancelled request stays CANCELLED even when its fetch
|
|
# failed: the finally block skips callbacks only for
|
|
# CANCELLED, and nobody is waiting on this fetch any more.
|
|
if request.status != FetchStatus.CANCELLED:
|
|
request.status = FetchStatus.FAILED
|
|
request.error = error_msg
|
|
|
|
result = FetchResult(
|
|
request_id=request.id,
|
|
success=False,
|
|
error=error_msg,
|
|
fetch_time=time.time() - start_time,
|
|
retry_count=request.retry_count
|
|
)
|
|
|
|
finally:
|
|
# Store result and clean up
|
|
with self._lock:
|
|
result.final_status = request.status
|
|
self.completed_requests[request.id] = result
|
|
if request.id in self.active_requests:
|
|
del self.active_requests[request.id]
|
|
# Stop accepting joiners and take the callback list in the same
|
|
# critical section. A submitter that arrives after this point
|
|
# finds no in-flight entry and either hits the cache (written
|
|
# above, before the result was built) or starts a fresh fetch --
|
|
# what it must never do is join a fetch whose callbacks have
|
|
# already run and then never be called.
|
|
if self._inflight_by_cache_key.get(request.cache_key) == request.id:
|
|
del self._inflight_by_cache_key[request.cache_key]
|
|
# A cancelled request delivers nothing: its joiners were told
|
|
# about a fetch that has been abandoned, and a replacement will
|
|
# call them via its own request.
|
|
if request.status == FetchStatus.CANCELLED:
|
|
callbacks = []
|
|
else:
|
|
callbacks = ([request.callback] if request.callback else [])
|
|
callbacks.extend(request.extra_callbacks)
|
|
|
|
# Update statistics
|
|
if result.success:
|
|
self.stats['completed_requests'] += 1
|
|
else:
|
|
self.stats['failed_requests'] += 1
|
|
|
|
self.stats['total_fetch_time'] += result.fetch_time
|
|
self.stats['average_fetch_time'] = (
|
|
self.stats['total_fetch_time'] /
|
|
(self.stats['completed_requests'] + self.stats['failed_requests'])
|
|
)
|
|
|
|
# Periodic cleanup after storing result
|
|
self._cleanup_completed_requests()
|
|
|
|
# Call every callback: the original submitter's and any that joined
|
|
# this fetch. One raising must not stop the others being delivered.
|
|
for cb in callbacks:
|
|
try:
|
|
cb(result)
|
|
except Exception as e:
|
|
logger.error(f"Error in callback for request {request.id}: {e}")
|
|
|
|
# Released after the loop, never inside it: every callback holds
|
|
# the same FetchResult (a sport's recent, upcoming and live
|
|
# managers usually share one fetch), so a release between
|
|
# deliveries would hand the later ones `result.data is None`.
|
|
#
|
|
# Only when there were callbacks: a request submitted without one
|
|
# collects its payload by polling get_result().
|
|
if callbacks:
|
|
self._release_payload(result)
|
|
request.result = None
|
|
|
|
return result
|
|
|
|
@staticmethod
|
|
def _release_payload(result: FetchResult) -> None:
|
|
"""Drop a delivered payload, keeping the result's status and timings.
|
|
|
|
Only called once EVERY callback has been handed the data -- callers
|
|
that joined an in-flight fetch share this object, so releasing between
|
|
deliveries strips the payload out from under the ones still queued.
|
|
Consumers read fetched data back from the cache under ``cache_key``;
|
|
the copy carried
|
|
here was pinning a parsed season schedule -- 946 games for NCAA
|
|
football, roughly a tenth of total RAM on a 1GB Pi -- in memory until
|
|
the hourly sweep.
|
|
|
|
The cache-hit path matters most: it runs once per update interval per
|
|
sport, mints a fresh request_id each time, and a memory-tier miss
|
|
re-parses the payload from disk. Those were genuinely separate copies
|
|
accumulating toward the 500-entry cap, not shared references.
|
|
"""
|
|
result.data = None
|
|
|
|
def _fetch_in_date_chunks(self, request: FetchRequest) -> Optional[Dict[str, Any]]:
|
|
"""Re-fetch a rejected ``YYYYMMDD-YYYYMMDD`` range as month/day chunks.
|
|
|
|
None means the request was not a day range, or every chunk failed; the
|
|
caller then re-raises the original 400 instead of caching an empty
|
|
season. See src/common/espn_dates.py.
|
|
"""
|
|
logger.info("Recovering %s %s from a rejected date range", request.sport, request.year)
|
|
return fetch_espn_date_chunks(
|
|
self.session,
|
|
request.url,
|
|
params=request.params,
|
|
headers=request.headers,
|
|
timeout=request.timeout,
|
|
logger=logger,
|
|
)
|
|
|
|
def _make_request_with_retry(self, request: FetchRequest) -> requests.Response:
|
|
"""
|
|
Make HTTP request with retry logic and exponential backoff.
|
|
|
|
Args:
|
|
request: Fetch request containing request details
|
|
|
|
Returns:
|
|
HTTP response
|
|
|
|
Raises:
|
|
requests.RequestException: If all retries fail
|
|
"""
|
|
last_exception = None
|
|
|
|
for attempt in range(request.max_retries + 1):
|
|
try:
|
|
response = self.session.get(
|
|
request.url,
|
|
params=request.params,
|
|
headers=request.headers,
|
|
timeout=request.timeout
|
|
)
|
|
return response
|
|
|
|
except requests.RequestException as e:
|
|
last_exception = e
|
|
request.retry_count = attempt + 1
|
|
|
|
if attempt < request.max_retries:
|
|
# Exponential backoff: 1s, 2s, 4s, 8s...
|
|
delay = 2 ** attempt
|
|
logger.warning(f"Request failed (attempt {attempt + 1}/{request.max_retries + 1}), retrying in {delay}s: {e}")
|
|
time.sleep(delay)
|
|
else:
|
|
logger.error(f"All {request.max_retries + 1} attempts failed for {request.sport} {request.year}")
|
|
|
|
raise last_exception
|
|
|
|
def get_result(self, request_id: str) -> Optional[FetchResult]:
|
|
"""
|
|
Get the result of a fetch request.
|
|
|
|
Args:
|
|
request_id: Request ID to get result for
|
|
|
|
Returns:
|
|
Fetch result if available, None otherwise
|
|
"""
|
|
# Periodic cleanup
|
|
self._cleanup_completed_requests()
|
|
|
|
with self._lock:
|
|
return self.completed_requests.get(request_id)
|
|
|
|
def is_request_complete(self, request_id: str) -> bool:
|
|
"""
|
|
Check if a request has completed.
|
|
|
|
Args:
|
|
request_id: Request ID to check
|
|
|
|
Returns:
|
|
True if request is complete, False otherwise
|
|
"""
|
|
# Periodic cleanup
|
|
self._cleanup_completed_requests()
|
|
|
|
with self._lock:
|
|
return request_id in self.completed_requests
|
|
|
|
def get_request_status(self, request_id: str) -> Optional[FetchStatus]:
|
|
"""
|
|
Get the status of a fetch request.
|
|
|
|
Args:
|
|
request_id: Request ID to get status for
|
|
|
|
Returns:
|
|
Request status if found, None otherwise
|
|
"""
|
|
with self._lock:
|
|
if request_id in self.active_requests:
|
|
return self.active_requests[request_id].status
|
|
elif request_id in self.completed_requests:
|
|
result = self.completed_requests[request_id]
|
|
if result.final_status is not None:
|
|
return result.final_status
|
|
return FetchStatus.COMPLETED if result.success else FetchStatus.FAILED
|
|
return None
|
|
|
|
def cancel_request(self, request_id: str) -> bool:
|
|
"""
|
|
Cancel a pending or in-progress request.
|
|
|
|
Args:
|
|
request_id: Request ID to cancel
|
|
|
|
Returns:
|
|
True if request was cancelled, False if not found or already complete
|
|
"""
|
|
with self._lock:
|
|
if request_id in self.active_requests:
|
|
request = self.active_requests[request_id]
|
|
if request.commit_claimed:
|
|
# Too late: the worker holds an authorised commit. Report
|
|
# the failure rather than half-cancelling a request whose
|
|
# data is about to land in the cache.
|
|
logger.debug(
|
|
"Not cancelling %s: its response is already being "
|
|
"committed", request_id
|
|
)
|
|
return False
|
|
request.status = FetchStatus.CANCELLED
|
|
del self.active_requests[request_id]
|
|
# Cancelling is the other way a request leaves active_requests,
|
|
# so the in-flight index has to be released here too or the key
|
|
# stays pointed at a request that no longer exists.
|
|
if self._inflight_by_cache_key.get(request.cache_key) == request_id:
|
|
del self._inflight_by_cache_key[request.cache_key]
|
|
logger.info(f"Cancelled request {request_id}")
|
|
return True
|
|
return False
|
|
|
|
def get_statistics(self) -> Dict[str, Any]:
|
|
"""
|
|
Get service statistics.
|
|
|
|
Returns:
|
|
Dictionary containing service statistics
|
|
"""
|
|
with self._lock:
|
|
return {
|
|
**self.stats,
|
|
'active_requests': len(self.active_requests),
|
|
'completed_requests_count': len(self.completed_requests),
|
|
'max_completed_requests': self._max_completed_requests,
|
|
'completed_requests_usage_percent': (len(self.completed_requests) / self._max_completed_requests * 100) if self._max_completed_requests > 0 else 0,
|
|
'last_cleanup': self._last_completed_requests_cleanup,
|
|
'cleanup_interval': self._completed_requests_cleanup_interval
|
|
}
|
|
|
|
def log_memory_stats(self):
|
|
"""Log current memory usage statistics."""
|
|
stats = self.get_statistics()
|
|
logger.info(f"BackgroundDataService Memory - Active: {stats['active_requests']}, "
|
|
f"Completed: {stats['completed_requests_count']}/{stats['max_completed_requests']} "
|
|
f"({stats['completed_requests_usage_percent']:.1f}%), "
|
|
f"Last cleanup: {time.time() - stats['last_cleanup']:.1f}s ago")
|
|
|
|
def _cleanup_completed_requests(self, force: bool = False) -> int:
|
|
"""
|
|
Automatically clean up old completed requests.
|
|
|
|
Args:
|
|
force: If True, perform cleanup regardless of time interval
|
|
|
|
Returns:
|
|
Number of requests removed
|
|
"""
|
|
now = time.time()
|
|
|
|
# Check if cleanup is needed
|
|
if not force and (now - self._last_completed_requests_cleanup) < self._completed_requests_cleanup_interval:
|
|
return 0
|
|
|
|
with self._lock:
|
|
removed_count = 0
|
|
current_time = time.time()
|
|
|
|
# Remove requests older than 1 hour
|
|
cutoff_time = current_time - 3600 # 1 hour
|
|
|
|
to_remove = []
|
|
for request_id, result in self.completed_requests.items():
|
|
# Check if request is old enough to remove
|
|
if result.completed_at < cutoff_time:
|
|
to_remove.append(request_id)
|
|
|
|
# Also enforce size limit if we have too many requests
|
|
if len(self.completed_requests) > self._max_completed_requests:
|
|
# Sort by completion time (oldest first)
|
|
sorted_requests = sorted(
|
|
self.completed_requests.items(),
|
|
key=lambda x: x[1].completed_at
|
|
)
|
|
|
|
# Remove oldest entries until we're under the limit
|
|
excess_count = len(self.completed_requests) - self._max_completed_requests
|
|
for i in range(excess_count):
|
|
if i < len(sorted_requests):
|
|
request_id = sorted_requests[i][0]
|
|
if request_id not in to_remove:
|
|
to_remove.append(request_id)
|
|
|
|
# Remove the requests
|
|
for request_id in to_remove:
|
|
del self.completed_requests[request_id]
|
|
removed_count += 1
|
|
|
|
self._last_completed_requests_cleanup = current_time
|
|
|
|
if removed_count > 0:
|
|
logger.debug(f"Cleaned up {removed_count} old completed requests (remaining: {len(self.completed_requests)})")
|
|
|
|
return removed_count
|
|
|
|
def shutdown(self, wait: bool = True):
|
|
"""
|
|
Shutdown the background data service.
|
|
|
|
Args:
|
|
wait: Whether to wait for active requests to complete
|
|
"""
|
|
logger.info("Shutting down BackgroundDataService...")
|
|
|
|
self._shutdown = True
|
|
|
|
# Cancel all active requests
|
|
with self._lock:
|
|
for request_id in list(self.active_requests.keys()):
|
|
self.cancel_request(request_id)
|
|
|
|
self.executor.shutdown(wait=wait)
|
|
|
|
logger.info("BackgroundDataService shutdown complete")
|
|
|
|
def __del__(self):
|
|
"""Cleanup when service is destroyed."""
|
|
if not self._shutdown:
|
|
self.shutdown(wait=False)
|
|
|
|
# Global service instance
|
|
_background_service: Optional[BackgroundDataService] = None
|
|
_service_lock = threading.Lock()
|
|
|
|
def get_background_service(cache_manager=None, max_workers: int = 3) -> BackgroundDataService:
|
|
"""
|
|
Get the global background data service instance.
|
|
|
|
Args:
|
|
cache_manager: Cache manager instance (required for first call)
|
|
max_workers: Maximum number of background threads
|
|
|
|
Returns:
|
|
Background data service instance
|
|
"""
|
|
global _background_service
|
|
|
|
with _service_lock:
|
|
if _background_service is None:
|
|
if cache_manager is None:
|
|
raise ValueError("cache_manager is required for first call to get_background_service")
|
|
_background_service = BackgroundDataService(cache_manager, max_workers)
|
|
|
|
return _background_service
|
|
|
|
def shutdown_background_service():
|
|
"""Shutdown the global background data service."""
|
|
global _background_service
|
|
|
|
with _service_lock:
|
|
if _background_service is not None:
|
|
_background_service.shutdown()
|
|
_background_service = None
|