feat(fetch): shared fetch service, stage 1 (pooling, merging, host budgets, counters) (#702)

Core's own HTTP fetch paths (APIHelper, fetch_espn_scoreboard and its date chunks, BackgroundDataService, BaseOddsManager.get_odds) go through one service in src/common/fetch_service.py: shared connection pools per retry policy, merged identical in-flight GETs, per-host token-bucket budgets (fetch_service.rate_limits), and per-plugin request counters published to GET /api/v3/plugins/fetch-stats. Return values, exceptions, cache keys, TTLs and retry policies are unchanged. Core-internal in this release; plugins should not import it directly yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-10-01 15:38:11 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent 9edeb6da14
commit 4be53d048b
19 changed files with 2407 additions and 31 deletions
+21 -1
View File
@@ -27,6 +27,7 @@ Rules for the package:
| [`bdf_font`](#bdf_font) | Load and draw BDF bitmap fonts | Yes, if drawing BDF text directly | 3.5.0 |
| [`espn_dates`](#espn_dates) | Fetch ESPN scoreboards across a date range | Yes (scoreboards) | 3.5.0 |
| [`favorite_team_check`](#favorite_team_check) | Log why a favourite team code shows nothing | Yes (scoreboards) | 3.6.0 |
| [`fetch_service`](#fetch_service) | Pooled, merged, budgeted and counted HTTP for core fetch paths | No, core-internal (reached through `api_helper` and `espn_dates`) | n/a |
| [`font_layout`](#font_layout) | Reproducible TrueType loading, crisp sizes | Yes | 3.4.0 |
| [`frame_timing`](#frame_timing) | Timing of every presented frame, stall watchdog | No, core-internal | n/a |
| [`json_body`](#json_body) | Parse a response body as JSON, with orjson if installed | Optional (large payloads) | 3.5.0 |
@@ -108,7 +109,9 @@ and truncates results when `limit` is above 500. `fetch_espn_scoreboard()`
splits a range into month and day requests ESPN accepts and merges the
results; `espn_date_chunks()`, `fetch_espn_date_chunks()`,
`clamp_espn_limit()` and `merge_scoreboard_payloads()` are the pieces.
Scoreboard plugins also bundle a copy for older cores.
Every request goes through [`fetch_service`](#fetch_service), the chunks
counted against the plugin that asked. Scoreboard plugins also bundle a copy
for older cores.
### favorite_team_check
@@ -121,6 +124,23 @@ says the league has nothing on yet; `reset()` re-arms it after a config edit.
Diagnostics only: every failure is swallowed. Scoreboard plugins also bundle
a copy for older cores.
### fetch_service
[`fetch_service.py`](fetch_service.py). Core-internal for now. Every core
fetch path -- `APIHelper.get`/`post`, `espn_dates` (so every scoreboard's
ESPN scoreboard fetch and `SportsFetchMixin`), `BackgroundDataService` and
`BaseOddsManager` -- calls `fetch_get(session, url, ...)` instead of
`session.get(url, ...)`. Same arguments, return value and exceptions; on top
it shares one connection pool per host per retry policy
(`share_connection_pool`), merges identical GETs in flight, applies per-host
token buckets (`fetch_service.rate_limits` in config.json; ESPN gets 20/s,
burst 200), revalidates with server-sent `ETag`/`Last-Modified` and counts
requests per plugin and per host. The display publishes the counters
(`FetchStatsPublisher`) for `GET /api/v3/plugins/fetch-stats`. Which plugin
made a request comes from `plugin_scope()`, set by the plugin executor, or
else from the plugin directory on the stack. See
[docs/PLUGIN_API_REFERENCE.md](../../docs/PLUGIN_API_REFERENCE.md#fetching-data).
### font_layout
[`font_layout.py`](font_layout.py). `load_truetype(path, size)` is
+14 -7
View File
@@ -11,10 +11,10 @@ import time
from datetime import datetime
from types import MappingProxyType
from src.common.espn_dates import ESPN_MAX_LIMIT
from src.common.fetch_service import fetch_get, fetch_post, share_connection_pool
from typing import TYPE_CHECKING, Any, Dict, Mapping, Optional, cast
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
if TYPE_CHECKING:
@@ -45,7 +45,11 @@ class APIHelper:
- Requests go through one ``requests.Session`` that retries GET, HEAD
and OPTIONS on 429 and 5xx with exponential backoff, and sends
:data:`DEFAULT_HTTP_HEADERS`.
:data:`DEFAULT_HTTP_HEADERS`. Its connection pool is shared with every
other helper using the same retry policy, and requests go through the
core fetch service (``src/common/fetch_service.py``): identical GETs in
flight are merged, hosts with a budget are paced, and requests are
counted per plugin. Return values and errors are unchanged.
- Consecutive requests from one helper are spaced at least
``set_rate_limit()`` seconds apart (1 second by default). A cache hit
does not count.
@@ -81,9 +85,10 @@ class APIHelper:
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "HEAD", "OPTIONS"]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
self.session.mount("https://", adapter)
self.session.mount("http://", adapter)
# The shared adapter for this retry policy: the same retries as a
# private HTTPAdapter(max_retries=retry_strategy), with the connection
# pool shared by every helper (fetch_service).
share_connection_pool(self.session, retry_strategy)
self.session.headers.update({**DEFAULT_HTTP_HEADERS, 'Connection': 'keep-alive'})
@@ -128,7 +133,8 @@ class APIHelper:
request_headers.update(headers)
# Make request
response = self.session.get(
response = fetch_get(
self.session,
url,
params=params,
headers=request_headers,
@@ -255,7 +261,8 @@ class APIHelper:
if headers:
request_headers.update(headers)
response = self.session.post(
response = fetch_post(
self.session,
url,
data=data,
json=json_data,
+29 -3
View File
@@ -32,6 +32,7 @@ scoreboards ask every 30 seconds. After that the range is tried again, so the
workaround retires itself if ESPN reverts.
"""
import contextvars
import threading
import time
from concurrent.futures import ThreadPoolExecutor
@@ -47,6 +48,21 @@ except ImportError:
def response_json(response: Any) -> Any:
return response.json()
try:
# The core fetch service: counts, per-host budget, merging of identical
# requests. Same call, same result and errors as ``session.get``.
from src.common.fetch_service import fetch_get, pinned_caller
except ImportError:
# Bundled copies on cores without it call the session directly.
import contextlib
def fetch_get(session: Any, url: str, *, share_in_flight: bool = True,
**kwargs: Any) -> Any:
return session.get(url, **kwargs)
def pinned_caller() -> Any:
return contextlib.nullcontext()
# Above this, ESPN returns a truncated list instead of an error. See module
# docstring: 500 is the largest value measured to return complete data.
ESPN_MAX_LIMIT = 500
@@ -195,7 +211,8 @@ def _fetch_one_chunk(
logged and swallowed here rather than raised to the gather below.
"""
try:
response = session.get(
response = fetch_get(
session,
url,
params=dict(params, dates=chunk, limit=ESPN_MAX_LIMIT),
headers=headers,
@@ -220,6 +237,10 @@ def _fetch_chunks(
callers keep ``chunks`` order from the returned list -- but it does mean
the session is shared across threads, which is why this only ever issues
GETs and never touches session state.
Each chunk runs in a copy of the caller's context, with the caller pinned
into it, so the fetch service counts the chunks against the plugin that
asked for the range rather than against the core.
"""
if not chunks:
return []
@@ -229,10 +250,15 @@ def _fetch_chunks(
if len(chunks) == 1:
return [fetch(chunks[0])]
workers = min(ESPN_CHUNK_WORKERS, len(chunks))
with pinned_caller():
# One copy per chunk: a Context cannot be entered by two threads.
contexts = [contextvars.copy_context() for _ in chunks]
with ThreadPoolExecutor(
max_workers=workers, thread_name_prefix="espn-chunk",
) as pool:
return list(pool.map(fetch, chunks))
futures = [pool.submit(context.run, fetch, chunk)
for context, chunk in zip(contexts, chunks)]
return [future.result() for future in futures]
def fetch_espn_date_chunks(
@@ -363,7 +389,7 @@ def fetch_espn_scoreboard(
# real error to log, without spending the chunks a second time.
chunks_tried = True
response = session.get(url, params=params, headers=headers, timeout=timeout)
response = fetch_get(session, url, params=params, headers=headers, timeout=timeout)
if is_range and response.status_code == 400 and not chunks_tried:
_note_range_rejected()
if logger:
File diff suppressed because it is too large Load Diff