Every instrument we have read line by line. A defect is named only where a public receipt exists. Everything else is counted, not claimed.
Every row here resolves to a live URL on someone else's repository.
Reset-scope and grading-path integrity
Grader-host integrity and timeout semantics
Grader parse and template integrity
Answer extraction, caching, math grading
Verdict extraction, error handling, judge prompts
Per-task grader integrity
Judge extractors and math scoring
Judge postprocess and evaluated-text execution
Verdict parser integrity
Judge JSON extraction
Judge parsing, scoring inputs, query integrity
Win-rate construction and parser steering
Leaderboard verification and anti-cheat integrity
Detector normalization and ASR accounting
Scorer semantics for could-not-score
Monitor prompt construction
Validation-metric aggregation
Scanner fail-open semantics
Reproduced defects in these exist in the private record and appear in the ledger without a name, either because disclosure is still on the vendor's clock or because the report is not filed yet. Listed, not claimed.
Gold = read adversarially and held
We read the scoring path, not the README. The grader, the parser, the aggregation, the cache, the reset.
Every candidate defect must reproduce against the genuine code path. Hypotheses that fail to execute are recorded as refuted and never reported.
Every finding is checked against the project's own tracker and the public literature before it is counted. Known issues are cited, not re-claimed.
Reports go to the maintainer before anywhere else, with the reproduction and a fix. Publication follows on the vendor's clock.
If it is, you can read what we found. If it is not, it can be.
You send one public claim. A benchmark score, a safety rate, a leaderboard position.
We read the instrument, not the number. Harness, judge, grading code.
One page back: what holds, what is unverifiable, what is broken. Each with a reproduction.