8.5 KiB
Audit Standard (canonical)
This is the fixed rubric. Every module audit must follow it verbatim so scores are comparable across modules, runs, and projects. Do not improvise scoring.
Scoring rubric (0–100) → grade
| Score | Grade | Meaning |
|---|---|---|
| 90–100 | A | Clean, well-scoped, idiomatic. No material smells. |
| 75–89 | B | Minor issues: a documented shim, mild bloat, a localized cast. |
| 60–74 | C | Notable hacks/fallbacks, real bloat, or duplication that has a clear owner. |
| 40–59 | D | Significant legacy/stubs/duplication, or a dual-format/protocol violation. |
| 0–39 | F | Broken, fake output, or an unfinished feature wired in as if done. |
Be rigorous and evidence-based, not generous. A module with one HIGH finding rarely scores above 60; with only LOW findings it usually scores 80+.
Smell taxonomy (the tags)
Use these exact tag strings. clean is the only positive tag; the rest are negative
(the map colors them red and counts them in the report).
monkeypatch— runtime mutation of another module / stdlib / vendor;setattron foreign objects;sys.modules/sys.meta_pathsurgery; reassigning store actions.fallback— "try the real thing, then fake/degrade"; chaineda || b || c/a ?? bdefaults that hide which value is real.silent-except/silent-catch—except: pass, bareexcept, emptycatch {}that swallow errors with no log/signal.legacy— deprecated/back-compat shims, retired vocabulary, dead-but-shipped code, parallel "old + new" code paths kept side by side.dual-format— accepting both snake_case and camelCase (or two payload shapes) for the same field; the classicdisplay_name || displayNamepatch.stub/placeholder—NotImplemented,TODO: implement, dead buttons, demo scripts, hardcoded sample data presented as real.fake-output— returns random/canned/hardcoded results where real computation is implied.duplication— logic copy-pasted from a sibling or that an existing shared abstraction already covers.bloat/god-component— oversized file/function; many responsibilities in one unit.glue— thin, valueless pass-through / boilerplate forwarding: rows of one-line wrappers that only forward args to another layer (e.g. dozens ofsend({type:...})methods), an adapter that copies a payload field-for-field without transforming, a store/function that only re-exports or delegates to another. The indirection earns nothing. Distinct frombloat(size) andduplication(copy-paste): glue is about forwarding that adds no value.any-escape—as any,@ts-ignore, untyped boundaries used to bypass the type system.over-fit— hardcoded to one case where a small generalization was expected.clean— no material issues.
Judgement rules:
- A documented, bounded compat shim that deliberately refuses to silently coerce is
legacyat most LOW — do not over-penalize disciplined shims. - A fallback that is a real security control or numeric guard (e.g. identity matrix on singular input, stripping untrusted shaders) is not a smell.
- Native-dependency gating that raises or returns an error when a lib is missing is
correct; only
fake-outputif it silently returns fabricated data. - A
*Placeholdername is not automatically a stub — read it; it may be a finished read-only widget. - A single thin delegator, or a genuine boundary normalizer that converts/validates
once, is fine — not
glue. Flagglueonly when pass-through wrappers proliferate (many near-identical forwarders that should be collapsed, generated, or replaced by a generic dispatch) or an adapter forwards with no transformation. Usually MED when it proliferates, LOW for a one-off.
Severity (each finding)
HIGH— wrong/dangerous/fake, a protocol or security issue, or a god-file that is a genuine maintenance hazard.MED— a real smell a maintainer should fix: a live dual-format patch, an un-migrated duplicate, an unfinished-but-wired path.LOW— a documented shim, a cosmetic cast, benign bloat. Worth noting, not urgent.
Always cite file:line and quote/paraphrase the offending snippet. Never report a
grep hit as a problem without reading the surrounding code.
Independent-subagent protocol (REQUIRED)
Every module's score MUST be produced by a separate subagent, never inline in the main thread, and never reused across modules. One subagent audits one module against the paths in its state entry. Spawn them in parallel (Explore or general-purpose).
Subagent prompt template
You are auditing CODE QUALITY of ONE functional module for an architecture audit. Module: {label} (
{id}). Files: {paths}. Project root: {root}.Read the module's code (grep the markers below, then READ the surrounding code — never flag a grep hit you haven't read). Judge it against this rubric: {paste the "Scoring rubric", "Smell taxonomy", "Severity" sections above}
Hunt specifically for: monkeypatch / stdlib mutation, fallback chains & silent excepts, legacy/deprecated/back-compat shims, dual-format (snake||camel) handling, stubs / fake output / unfinished-but-wired code, bloat/god-files, duplication of logic that exists elsewhere, thin valueless glue (proliferating pass-through wrappers / no-op adapters), and over-fitting. Also state whether the module is appropriately generic.
Return ONLY this JSON (no prose): {"score": <0-100>, "grade": "<A|B|C|D|F>", "tags": ["", ...], "findings": [{"sev":"HIGH|MED|LOW","loc":"file:line","text":"concrete issue + evidence"}, ...]} If clean, use tags ["clean"] and findings []. Be rigorous, not generous.
Use schema on the Agent call to force that JSON shape when available. Then feed each
result to scripts/apply_audit.py --id <id> --json '<result>'.
Test-author protocol (the test command + the baseline step of fix)
A dedicated test-author subagent generates tests for a module. It is separate from the auditor, fixer, and verifier.
- Detect, don't invent. Find the repo's test framework and location from existing tests near the module (pytest / jest / vitest / go test / …); match their style and placement. Never introduce a new framework or harness.
- characterization mode (default before a fix): capture the module's CURRENT observable behavior — inputs→outputs, side effects, payload shapes — as assertions of "same as today", not "correct". Target the public surface; don't pin private internals. Use snapshot/golden tests only where the repo already does.
- coverage mode: cover the public API and the specific behaviors named in the
module's
findings. Aim for meaningful branches, not line count. - Must be GREEN on the current, unmodified code before returning. If a test you want to write fails because of a real bug, FLAG it as a finding — do not assert the buggy output as if it were the desired behavior.
- Tests are real, committed source (the regression net); never delete or weaken them to move a number.
- Return JSON:
{"framework":"...","files":["..."],"locked":"<behaviors locked>", "gaps":"<what is still uncovered>","flagged":[{"sev":"...","loc":"...","text":"..."}]}.
Acceptance / regression gate (the gate in fix)
An acceptance/verifier subagent, independent of the fixer, proves a fix introduced no regression. It does NOT score quality — pass/fail only.
- Re-run the EXACT baseline test set captured before the fix (same commands), plus the narrowest build/typecheck for the touched area.
- PASS iff every test that was green before is green after, AND no new build / type / lint errors appeared. A baseline-green test that is now failing, errored, skipped, deleted, or flaky counts as a regression → FAIL (you cannot remove a test to pass).
- New tests the fixer may have added are ignored (and the fixer should not add any).
- Return JSON:
{"pass": <bool>, "ran": "<commands>", "regressions": [{"test":"...","was":"pass","now":"fail|error|missing","evidence":"..."}], "evidence": "<short summary>"}. - Gate rule: no PASS → no re-audit and no rendered score improvement. Report the failure with evidence; revert or hand back to the fixer.
Cross-cutting themes
After all modules are scored, the orchestrator (main thread) writes 4–7
reportThemes into modules.json — patterns seen across modules (e.g. "dual-format
recurs in N handlers", "duplication between X and Y", "stub backend wired live"). Each
is [headline, body]. These are synthesis, not per-module scoring, so the main thread
writes them.