mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-08-01 16:58:06 +00:00
* chore(ci): add security-audit workflow and plugin security-proof scripts - scripts/prove_security.py, audit_plugins.py, generate_report.py -- automated checks (dangerous eval()/exec() calls, dependency scanning, report generation) for plugins. - .github/workflows/security-audit.yml + bandit.yaml -- CI wiring for the above plus gitleaks secret scanning and bandit static analysis. - .github/workflows/tests.yml -- pytest matrix across Python 3.10-3.12. Also fixes two Codacy findings while these files are freshly landing: - prove_security.py: dropped a pointless f-string prefix with no placeholders. - security-audit.yml: pinned gitleaks/gitleaks-action to a full commit SHA (matching this repo's existing pinning convention in test.yml) instead of the floating v2 tag. Split out of the original chore/dead-code-removal commit, which had accidentally bundled this in alongside unrelated dead-code deletions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * chore: drop workflow files -- pushed separately (needs workflow OAuth scope) * fix(security-tooling): address PR review findings across bandit.yaml, audit_plugins.py, generate_report.py, prove_security.py bandit.yaml: - Removed scripts/prove_security.py's file-level exclusion. Ran bandit directly to get ground truth: the real false positive is B105 (dict key "PASS" misread as password-like), not the eval/exec pattern the old comment claimed. Added a targeted # nosec B105 there, and found+fixed the identical pattern already present in generate_report.py. - Left the repo-wide B607 skip as-is: confirmed via AST scan that properly narrowing it touches 100+ bare-name subprocess call sites across wifi_manager.py, store_manager.py, permission_utils.py, app.py, and start.py -- none of which are part of this PR. Out of proportion to fix here; flagged as a dedicated follow-up. scripts/audit_plugins.py: - SyntaxError/OSError while scanning a plugin file now report CRITICAL (blocking) instead of WARNING/INFO -- a file that couldn't be parsed or read was never actually checked for danger, so it must not silently pass the audit. - --plugin <name> now tracks whether the requested plugin was found across all PLUGIN_BASE_DIRS and exits 1 with a clear error if not, instead of silently scanning zero plugins and reporting success. - The AST visitor now tracks import aliases (import subprocess as sp; from builtins import eval as e) and resolves them before checking against dangerous APIs, closing a straightforward evasion of every PLUGIN-001 through PLUGIN-005 check. Verified against both aliased and unaliased evasion patterns. scripts/generate_report.py: - _md_table_row now escapes pipe characters and normalizes newlines in every cell, so scanner-controlled content (a matched secret, a bandit issue_text) can't corrupt the Markdown table structure. - _load now distinguishes "artifact missing/malformed" from "valid empty result": each summarizer returns an availability flag, and main() now reports INCOMPLETE (not PASSED) with exit code 1 when any artifact is unavailable, instead of silently folding it in as 0 findings. - Gitleaks suppression now uses exact-match placeholder values (pulled from the actual config_secrets.template.json) plus a template-path allowlist, replacing broad substring checks that could hide a real secret containing something like "example.com" as part of its value. scripts/prove_security.py: - T1b (dangerous plugin calls): a file that fails to parse/read now reports CRITICAL with the exception details instead of being silently swallowed by `except (SyntaxError, OSError): pass`. - T6 (Docker hardening): base images must now be pinned to an @sha256 digest; a specific tag like python:3.12 is mutable and is now correctly flagged as unpinned, not just missing tags or :latest. - T2a (API surface): no config mechanism for enforcing local-only access exists in this codebase today (app.py hardcodes host='0.0.0.0'), so the "environment-aware" check as described isn't buildable without adding new config infrastructure -- out of scope here. Applied the achievable part: upgraded from INFO to WARNING, since enforcement can never currently be confirmed. - T1a (zip-slip): replaced the whole-file substring check with an AST walk that finds every extract()/extractall() call and confirms an is_relative_to() guard + "Zip-slip detected" log precede it in the same function. Verified it still passes on the real store_manager.py (both the per-member and validate-then-bulk-extract call sites) and correctly flags a synthetic unguarded extractall(). - T3a (hardcoded secrets): violation details no longer include the matched credential text -- only file, line, pattern type, and a redacted SHA-256 fingerprint, so a real finding doesn't get published into CI logs/artifacts/PR comments with wider exposure than the original leak. Verified with a synthetic secret that no raw content reaches the output. Validated: all four files compile; bandit scans all three scripts clean (2 legitimate targeted suppressions, 0 unaddressed findings); each new/ changed code path exercised directly (alias evasion, unmatched --plugin, missing/malformed/valid-empty artifacts, digest-pinning, zip-slip guard/no-guard, secret redaction); full audit_plugins.py -> prove_security.py -> generate_report.py pipeline run end-to-end producing a correct report. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * fix(security-tooling): address follow-up review findings on PR #414 scripts/generate_report.py: - Added an explicit `object` type annotation to _md_sanitize_cell's value parameter -- it deliberately accepts any stringifiable value (calls str(value) unconditionally), so `object` reflects its actual contract more accurately than leaving it untyped. scripts/prove_security.py: - Dockerfile FROM-line parsing: renamed the comprehension variable `l` to `line` (ambiguous single-letter name). More importantly, fixed a real false-positive: `FROM --platform=<platform> <image>` was reading the --platform= flag itself as the image token, so a properly digest-pinned image behind a platform flag was incorrectly reported as unpinned. Verified against platform+digest, platform+tag-only, and digest+AS-alias Dockerfiles. scripts/audit_plugins.py: - Consolidated visit_Call's dangerous-API detection: previously, alias resolution only covered ast.Name calls for eval/exec/compile and ast.Attribute calls for subprocess/os.system, missing from-imported subprocess/os functions called as bare names (from subprocess import run as prun; prun(cmd, shell=True) or from os import system as s; s(cmd)). Added _resolve_call_target() to resolve both call shapes to a single fully-qualified target, then run all five PLUGIN-00x checks against that one resolved value. Verified against 10 evasion combinations (from-import aliases, direct/attribute calls, aliased module imports) and confirmed zero false positives on benign os/ subprocess usage without shell=True. Validated: all three files compile, bandit scans clean (same 2 legitimate suppressions as before, 0 new findings), audit_plugins.py/prove_security.py re-run against the real repo with no regressions from the prior fix pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
357 lines
14 KiB
Python
357 lines
14 KiB
Python
#!/usr/bin/env python3
|
||
"""
|
||
Security Report Generator
|
||
|
||
Aggregates JSON output from all CI security audit jobs into a single
|
||
Markdown report suitable for PR comments and artifact storage.
|
||
|
||
Expected artifact layout (from actions/download-artifact@v4):
|
||
<artifact-dir>/
|
||
sast-results/
|
||
bandit-results.json
|
||
semgrep-results.json
|
||
dependency-audit-results/
|
||
pip-audit-results.json
|
||
safety-results.json
|
||
secrets-scan-results/
|
||
gitleaks-results.json
|
||
security-proofs-results/
|
||
security-proofs-results.json
|
||
plugin-audit-results/
|
||
plugin-audit-results.json
|
||
|
||
Usage:
|
||
python scripts/generate_report.py --artifact-dir audit-artifacts/ --output report.md
|
||
python scripts/generate_report.py --artifact-dir audit-artifacts/ --output report.md --verbose
|
||
"""
|
||
|
||
import argparse
|
||
import json
|
||
import sys
|
||
from pathlib import Path
|
||
from datetime import datetime, timezone
|
||
|
||
PROJECT_ROOT = Path(__file__).resolve().parent.parent
|
||
|
||
# Gitleaks matches exactly equal to one of these (not a substring match -- a
|
||
# real secret that merely contains one of these words as part of its actual
|
||
# value must still be reported) are known template placeholders.
|
||
_GITLEAKS_SUPPRESS_EXACT_VALUES = {
|
||
"YOUR_YOUTUBE_API_KEY",
|
||
"YOUR_YOUTUBE_CHANNEL_ID",
|
||
"YOUR_GITHUB_PERSONAL_ACCESS_TOKEN",
|
||
}
|
||
|
||
# Findings in these files are suppressed regardless of value -- they are
|
||
# template/example files that are expected to only ever contain placeholders.
|
||
_GITLEAKS_SUPPRESS_PATHS = [
|
||
"config_secrets.template.json",
|
||
"config.template.json",
|
||
]
|
||
|
||
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
# Helpers
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
|
||
def _load(path: Path) -> tuple[dict | list | None, str | None]:
|
||
"""Load a JSON artifact file.
|
||
|
||
Returns (data, error): error is None on success (data is whatever was
|
||
parsed, which may legitimately be an empty list/dict for a clean scan);
|
||
otherwise error is a human-readable reason the artifact is unavailable,
|
||
distinguishing "missing/malformed artifact" from "valid empty result" so
|
||
callers don't silently treat a broken CI job as a clean pass.
|
||
"""
|
||
if not path.exists():
|
||
return None, f"artifact not found: {path}"
|
||
try:
|
||
return json.loads(path.read_text(encoding="utf-8")), None
|
||
except (json.JSONDecodeError, OSError) as exc:
|
||
return None, f"could not read/parse {path}: {exc}"
|
||
|
||
|
||
def _md_sanitize_cell(value: object) -> str:
|
||
"""Escape/normalize a value so scanner-controlled content (a matched
|
||
secret, a bandit issue_text, a file path) can't alter the Markdown
|
||
table's structure: pipes would add bogus columns, newlines would break
|
||
out of the row (or forge a fake header/separator line)."""
|
||
text = str(value)
|
||
text = text.replace("\\", "\\\\").replace("|", "\\|")
|
||
text = text.replace("\r\n", " ").replace("\n", " ").replace("\r", " ")
|
||
return text
|
||
|
||
|
||
def _md_table_row(*cells: str) -> str:
|
||
return "| " + " | ".join(_md_sanitize_cell(c) for c in cells) + " |"
|
||
|
||
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
# Per-tool summarizers
|
||
# Returns: (markdown_lines: list[str], critical_count: int, available: bool)
|
||
# `available=False` means the artifact was missing or malformed -- distinct
|
||
# from a valid scan that simply found nothing -- so the caller can report
|
||
# INCOMPLETE instead of silently counting it as a clean pass.
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
|
||
def _summarize_bandit(artifact_dir: Path) -> tuple[list[str], int, bool]:
|
||
data, error = _load(artifact_dir / "sast-results" / "bandit-results.json")
|
||
if error:
|
||
return [f"_bandit results unavailable: {error}_"], 0, False
|
||
|
||
results = data.get("results", [])
|
||
high = [r for r in results if r.get("issue_severity") == "HIGH"]
|
||
medium = [r for r in results if r.get("issue_severity") == "MEDIUM"]
|
||
low = [r for r in results if r.get("issue_severity") == "LOW"]
|
||
|
||
lines = [
|
||
f"**Bandit**: {len(high)} HIGH · {len(medium)} MEDIUM · {len(low)} LOW"
|
||
]
|
||
|
||
if high:
|
||
lines += [
|
||
"",
|
||
"| Severity | File | Line | Issue |",
|
||
"| --- | --- | --- | --- |",
|
||
]
|
||
for r in high[:10]:
|
||
fname = Path(r.get("filename", "")).name
|
||
lines.append(_md_table_row(
|
||
"HIGH", f"`{fname}`",
|
||
str(r.get("line_number", "?")),
|
||
r.get("issue_text", "")
|
||
))
|
||
if len(high) > 10:
|
||
lines.append(f"_… and {len(high) - 10} more HIGH findings_")
|
||
|
||
return lines, len(high), True
|
||
|
||
|
||
def _summarize_pip_audit(artifact_dir: Path) -> tuple[list[str], int, bool]:
|
||
data, error = _load(artifact_dir / "dependency-audit-results" / "pip-audit-results.json")
|
||
if error:
|
||
return [f"_pip-audit results unavailable: {error}_"], 0, False
|
||
|
||
# pip-audit JSON format: {"dependencies": [{"name": ..., "vulns": [...]}]}
|
||
vulns: list[dict] = []
|
||
for dep in data.get("dependencies", []):
|
||
for v in dep.get("vulns", []):
|
||
vulns.append({"package": dep.get("name", "?"), **v})
|
||
|
||
lines = [f"**pip-audit**: {len(vulns)} vulnerabilities found"]
|
||
|
||
if vulns:
|
||
lines += ["", "| Package | ID | Fix |", "| --- | --- | --- |"]
|
||
for v in vulns[:10]:
|
||
fix = v.get("fix_versions", ["none"])
|
||
fix_str = ", ".join(fix) if fix else "none"
|
||
lines.append(_md_table_row(
|
||
v.get("package", "?"),
|
||
v.get("id", "?"),
|
||
fix_str,
|
||
))
|
||
|
||
# Treat known vulnerabilities as warnings, not critical (they may be unavoidable)
|
||
return lines, 0, True
|
||
|
||
|
||
def _summarize_gitleaks(artifact_dir: Path) -> tuple[list[str], int, bool]:
|
||
data, error = _load(artifact_dir / "secrets-scan-results" / "gitleaks-results.json")
|
||
if error:
|
||
return [f"_gitleaks results unavailable: {error}_"], 0, False
|
||
|
||
if not isinstance(data, list):
|
||
data = []
|
||
|
||
real_findings = []
|
||
suppressed = 0
|
||
for finding in data:
|
||
secret_val = str(finding.get("Secret", "") or finding.get("Match", ""))
|
||
file_name = Path(finding.get("File", "")).name
|
||
if (secret_val in _GITLEAKS_SUPPRESS_EXACT_VALUES
|
||
or file_name in _GITLEAKS_SUPPRESS_PATHS):
|
||
suppressed += 1
|
||
else:
|
||
real_findings.append(finding)
|
||
|
||
lines = [
|
||
f"**Gitleaks**: {len(real_findings)} finding(s) "
|
||
f"({suppressed} suppressed as template placeholders)"
|
||
]
|
||
|
||
if real_findings:
|
||
lines += ["", "| Rule | File | Line | Description |", "| --- | --- | --- | --- |"]
|
||
for f in real_findings[:10]:
|
||
fname = Path(f.get("File", "")).name
|
||
lines.append(_md_table_row(
|
||
f.get("RuleID", "?"),
|
||
f"`{fname}`",
|
||
str(f.get("StartLine", "?")),
|
||
f.get("Description", ""),
|
||
))
|
||
|
||
critical = len(real_findings) # any real secret is critical
|
||
return lines, critical, True
|
||
|
||
|
||
def _summarize_security_proofs(artifact_dir: Path) -> tuple[list[str], int, bool]:
|
||
data, error = _load(artifact_dir / "security-proofs-results" / "security-proofs-results.json")
|
||
if error:
|
||
return [f"_security proofs results unavailable: {error}_"], 0, False
|
||
|
||
if not isinstance(data, list):
|
||
data = []
|
||
|
||
critical = [r for r in data if r.get("severity") == "CRITICAL"]
|
||
warnings = [r for r in data if r.get("severity") == "WARNING"]
|
||
passed = [r for r in data if r.get("severity") == "PASS"]
|
||
skipped = [r for r in data if r.get("severity") == "SKIP"]
|
||
|
||
lines = [
|
||
f"**Security Proofs**: "
|
||
f"{len(passed)} PASS · {len(warnings)} WARN · "
|
||
f"{len(critical)} CRITICAL · {len(skipped)} SKIP",
|
||
"",
|
||
]
|
||
|
||
_icon = {"PASS": "✅", "INFO": "ℹ️", "WARNING": "⚠️", # nosec B105 - severity labels, not credentials
|
||
"CRITICAL": "🚨", "SKIP": "⏭️"}
|
||
for r in data:
|
||
icon = _icon.get(r.get("severity", ""), "❓")
|
||
lines.append(
|
||
f"- {icon} **{r.get('test_id', '?')}**: {r.get('message', '')}"
|
||
)
|
||
if r.get("details") and r.get("severity") in ("CRITICAL", "WARNING"):
|
||
lines.append(f" - _{r['details']}_")
|
||
|
||
return lines, len(critical), True
|
||
|
||
|
||
def _summarize_plugin_audit(artifact_dir: Path) -> tuple[list[str], int, bool]:
|
||
data, error = _load(artifact_dir / "plugin-audit-results" / "plugin-audit-results.json")
|
||
if error:
|
||
return [f"_plugin audit results unavailable: {error}_"], 0, False
|
||
|
||
summary = data.get("summary", {})
|
||
findings = data.get("findings", [])
|
||
critical_findings = [f for f in findings if f.get("severity") == "CRITICAL"]
|
||
warning_findings = [f for f in findings if f.get("severity") == "WARNING"]
|
||
|
||
lines = [
|
||
f"**Plugin Audit**: {data.get('plugins_scanned', '?')} plugins scanned — "
|
||
f"{summary.get('critical', 0)} CRITICAL · {summary.get('warnings', 0)} WARNINGS"
|
||
]
|
||
|
||
if critical_findings:
|
||
lines += ["", "| Plugin | File | Line | Rule | Message |",
|
||
"| --- | --- | --- | --- | --- |"]
|
||
for f in critical_findings[:10]:
|
||
fname = Path(f.get("file", "")).name
|
||
lines.append(_md_table_row(
|
||
f.get("plugin_id", "?"),
|
||
f"`{fname}`",
|
||
str(f.get("line", "?")),
|
||
f.get("rule", "?"),
|
||
f.get("message", ""),
|
||
))
|
||
|
||
if warning_findings and not critical_findings:
|
||
lines.append(f"\n_{len(warning_findings)} warning(s) found — see artifact for details_")
|
||
|
||
return lines, summary.get("critical", 0), True
|
||
|
||
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
# Main
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
|
||
def main() -> int:
|
||
parser = argparse.ArgumentParser(
|
||
description="Generate consolidated security audit report",
|
||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||
)
|
||
parser.add_argument("--artifact-dir", required=True,
|
||
help="Directory containing downloaded CI artifacts")
|
||
parser.add_argument("--output", "-o", required=True,
|
||
help="Output Markdown file path")
|
||
parser.add_argument("--verbose", "-v", action="store_true")
|
||
args = parser.parse_args()
|
||
|
||
artifact_dir = Path(args.artifact_dir)
|
||
timestamp = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC")
|
||
|
||
bandit_lines, bandit_crit, bandit_ok = _summarize_bandit(artifact_dir)
|
||
pip_audit_lines, pip_audit_crit, pip_audit_ok = _summarize_pip_audit(artifact_dir)
|
||
gitleaks_lines, gitleaks_crit, gitleaks_ok = _summarize_gitleaks(artifact_dir)
|
||
proofs_lines, proofs_crit, proofs_ok = _summarize_security_proofs(artifact_dir)
|
||
plugins_lines, plugins_crit, plugins_ok = _summarize_plugin_audit(artifact_dir)
|
||
|
||
unavailable_tools = [
|
||
name for name, ok in [
|
||
("bandit", bandit_ok), ("pip-audit", pip_audit_ok),
|
||
("gitleaks", gitleaks_ok), ("security-proofs", proofs_ok),
|
||
("plugin-audit", plugins_ok),
|
||
] if not ok
|
||
]
|
||
|
||
total_critical = bandit_crit + pip_audit_crit + gitleaks_crit + proofs_crit + plugins_crit
|
||
if unavailable_tools:
|
||
# A missing/malformed artifact means that tool's checks never
|
||
# actually ran -- this must not be reported as a clean PASS just
|
||
# because the *artifacts that did load* found nothing.
|
||
overall = "INCOMPLETE ⚠️"
|
||
elif total_critical > 0:
|
||
overall = "ACTION REQUIRED 🚨"
|
||
else:
|
||
overall = "PASSED ✅"
|
||
|
||
def section(title: str, lines: list[str]) -> str:
|
||
return f"### {title}\n\n" + "\n".join(lines) + "\n"
|
||
|
||
incomplete_note = (
|
||
f"\n_⚠️ Incomplete: results unavailable for {', '.join(unavailable_tools)} "
|
||
f"— see the corresponding section(s) below for details_\n"
|
||
if unavailable_tools else ""
|
||
)
|
||
|
||
report = f"""## 🔒 Security Audit — {overall}
|
||
|
||
_Generated: {timestamp}_
|
||
{incomplete_note}
|
||
| Critical | High/Warn | Overall |
|
||
| :---: | :---: | :---: |
|
||
| {'🚨 ' + str(total_critical) if total_critical else '✅ 0'} | ⚠️ see below | {overall} |
|
||
|
||
---
|
||
|
||
{section('SAST — Bandit', bandit_lines)}
|
||
{section('Dependencies — pip-audit', pip_audit_lines)}
|
||
{section('Secrets — Gitleaks', gitleaks_lines)}
|
||
{section('LEDMatrix Security Proofs', proofs_lines)}
|
||
{section('Plugin Security Audit', plugins_lines)}
|
||
---
|
||
|
||
_Total critical findings: **{total_critical}**_
|
||
"""
|
||
|
||
output_path = Path(args.output)
|
||
output_path.write_text(report, encoding="utf-8")
|
||
|
||
if args.verbose:
|
||
print(f" Report written to: {output_path}")
|
||
print(f" Status: {overall}")
|
||
print(f" Critical findings: {total_critical}")
|
||
print(f" bandit={bandit_crit} pip-audit={pip_audit_crit} "
|
||
f"gitleaks={gitleaks_crit} proofs={proofs_crit} plugins={plugins_crit}")
|
||
if unavailable_tools:
|
||
print(f" Unavailable: {', '.join(unavailable_tools)}")
|
||
|
||
if unavailable_tools:
|
||
return 1
|
||
|
||
return 0
|
||
|
||
|
||
if __name__ == "__main__":
|
||
sys.exit(main())
|