Files
LEDMatrix/test/test_redaction.py
T
ChuckandClaude Opus 5.5 82f3a3a3e4 fix(redaction): make credential redaction linear, not quadratic (#631)
* fix(redaction): make URL-userinfo redaction linear, not quadratic

_REDACT_URL_USERINFO could start a match at every letter of a run of
scheme characters, and each attempt read to the end of the run looking
for `://`. On a long unbroken run of letters or digits (a hex digest, an
ID, part of a response body) that is quadratic: 1.6s for 20k characters.

The display service redacts every message, stack trace and context value
it publishes in the error snapshot, holding the aggregator lock, and
re.sub holds the GIL for the whole call, so one such exception stalled
every thread, render loop included (~0.5s measured for 20k chars of hex).
It also made test_snapshot_stays_small the slowest test in the suite by
far: 142s of a 383s run, 139s of it in this one regex.

A match may now only start where a run of scheme characters starts
(negative lookbehind). Leading digits and `+.-` are captured in group 1
so the substitution restores them, and the scheme still has to start
with a letter, so what gets redacted is unchanged: old and new output
were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k
~6ms, and test_snapshot_stays_small takes 0.8s.

test/test_redaction.py pins the exact output for schemes that begin after
digits or `+.-`, and bounds 50k-character runs at 1s; against the old
pattern those timing tests fail at 3-11s each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK

* fix(redaction): make Authorization-header redaction linear too

_REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two
`\s*` separated only by an optional quote. With no quote, a whitespace
run could be split between them in every possible way, and when no
credential followed (end of text, or `,` `"` `<` ...) the engine tried
them all before giving up: quadratic, 8s for `authorization:` and 20k
spaces, 17s with `Proxy-Authorization:` (tried again at the inner
`authorization`). Same stall as the URL pattern: re.sub holds the GIL,
and the display service redacts everything it publishes.

The quote and the whitespace after it are now one optional unit,
`\s*(?:["\']\s*)?`, which matches the same strings with only one way to
split them. Output is identical to the old pattern on 300k fuzzed
inputs; 20k spaces now take ~1.6ms. A scan of all three redaction
patterns over prefix/run/suffix shapes finds none left that scales
superlinearly.

test/test_redaction.py pins exact output for quoted, tabbed, multi-line
and credential-less headers, and bounds header + 20k whitespace at 1s;
against the previous pattern those fail at 8-17s each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:49:51 -04:00

111 lines
4.8 KiB
Python

"""redact_credentials must stay linear in the length of its input.
Regressions under test, both quadratic regexes in src/redaction.py:
- The URL-userinfo pattern (`scheme://user:password@`) could start a match at
every letter of a run of scheme characters, and each attempt read to the end
of the run looking for `://`: 1.6s for a 20k-character run.
- The Authorization-header pattern had two `\\s*` separated only by an
optional quote, so a header followed by whitespace and no credential tried
every split of that whitespace between them: 8s for 20k spaces.
The display service redacts every message, stack trace and context value it
publishes in the error snapshot, and re.sub holds the GIL throughout, so an
exception quoting a hex digest or a long ID stalled the render loop with it.
test_error_snapshot_cross_process.py's snapshot-size test spent 140s here.
The fixed patterns have to redact exactly what the old ones did.
"""
import time
import pytest
from src.redaction import redact_credentials
# Each timed input took seconds before the fix and takes about a millisecond
# after it; the bound leaves CI plenty of headroom while still failing on a
# quadratic pattern.
_TIME_LIMIT = 1.0
def _timed(text):
start = time.perf_counter()
result = redact_credentials(text)
return result, time.perf_counter() - start
class TestUrlUserinfo:
@pytest.mark.parametrize("text,expected", [
("401 for https://user:hunter2@example.com/api",
"401 for https://user:<redacted>@example.com/api"),
("HTTPS://USER:HUNTER2@EXAMPLE.COM",
"HTTPS://USER:<redacted>@EXAMPLE.COM"),
("git+ssh://deploy:hunter2@host/repo",
"git+ssh://deploy:<redacted>@host/repo"),
# The scheme starts after digits or +.- in the same run. Those
# characters must survive, and the password must still go.
("1http://user:hunter2@host", "1http://user:<redacted>@host"),
("+.-http://user:hunter2@host", "+.-http://user:<redacted>@host"),
("a1+http://user:hunter2@host", "a1+http://user:<redacted>@host"),
("see a://u:first@b and c://v:second@d",
"see a://u:<redacted>@b and c://v:<redacted>@d"),
])
def test_password_is_redacted_and_the_rest_kept(self, text, expected):
assert redact_credentials(text) == expected
def test_a_url_without_a_password_is_untouched(self):
text = "GET https://user@example.com/path failed"
assert redact_credentials(text) == text
class TestAuthorizationHeader:
@pytest.mark.parametrize("text,expected", [
("Authorization: Bearer eyJ.SECRET.sig", "Authorization: Bearer <redacted>"),
("Proxy-Authorization: Basic dXNlcg==", "Proxy-Authorization: Basic <redacted>"),
("authorization: barecredential", "authorization: <redacted>"),
# Whitespace and an opening quote around the value, in either order.
('authorization=" Bearer tok"', 'authorization=" Bearer <redacted>"'),
("authorization: ' tok'", "authorization: ' <redacted>'"),
("authorization:\n\tBearer tok", "authorization:\n\tBearer <redacted>"),
])
def test_credential_is_redacted_and_the_rest_kept(self, text, expected):
assert redact_credentials(text) == expected
@pytest.mark.parametrize("text", ["authorization: ", "authorization: , next"])
def test_a_header_without_a_credential_is_untouched(self, text):
assert redact_credentials(text) == text
class TestLinearTime:
@pytest.mark.parametrize("unit", ["x", "0123456789abcdef", "1a", "a+", "1"])
def test_long_scheme_character_runs(self, unit):
text = (unit * 50_000)[:50_000]
result, elapsed = _timed(text)
assert result == text
assert elapsed < _TIME_LIMIT, f"{elapsed:.2f}s to redact {len(text)} chars of {unit!r}"
def test_a_credential_after_a_long_run_is_still_found(self):
run = "ab12" * 10_000
result, elapsed = _timed(f"{run} https://user:hunter2@example.com")
assert result == f"{run} https://user:<redacted>@example.com"
assert elapsed < _TIME_LIMIT
@pytest.mark.parametrize("header,whitespace", [
("authorization:", " "),
("Proxy-Authorization:", "\t"),
("authorization=", "\n"),
])
def test_a_header_followed_by_long_whitespace(self, header, whitespace):
text = header + whitespace * 20_000 + ","
result, elapsed = _timed(text)
assert result == text
assert elapsed < _TIME_LIMIT, (
f"{elapsed:.2f}s to redact {header!r} and {len(text) - len(header)} more chars")
def test_a_credential_after_long_whitespace_is_still_found(self):
gap = " " * 20_000
result, elapsed = _timed(f"authorization:{gap}Bearer tok")
assert result == f"authorization:{gap}Bearer <redacted>"
assert elapsed < _TIME_LIMIT