The Map
The instruments we have read, what we read them for, and where each one stands. A finding is named only where a public receipt exists; everything else carries a status label, not a mechanism.
How we read
We read the scoring path, not the README. The grader, the parser, the aggregation, the cache, the reset.
Every candidate defect must reproduce against the genuine code path. Hypotheses that fail to execute are recorded as refuted and never reported.
Every finding is checked against the project's own tracker and the public literature before it is counted. Known issues are cited, not re-claimed.
Reports go to the maintainer before anywhere else, with the reproduction and a fix. Publication follows on the vendor's clock.
Benchmarks
Submission and denominator integrity
Reward pipeline and task-suite integrity
Evaluator crash paths and aggregation
Scorer quirks and repository hygiene
Competition guard and scorecard integrity
Metric correctness under multi-answer scoring
Labels re-verified; headline reproduces exactly
Coding grader residency and judgment integrity
String-matching and denominator semantics
Protected-scoring integrity across all seven task families
Task-family scorer integrity census
Successor-suite reward pipeline and user-simulator integrity
Headline metric aggregation under judge failure
Vendored harness grader and rubric-template integrity
Grader execution and equivalence semantics
Pass@k and executor grading integrity
Retrieval-task ground-truth integrity
Result artifact integrity
Curation-judge and data-pipeline integrity
Multitenant evaluation suite integrity
Harnesses and scoring libraries
Annotator execution and score extraction
Judge parsing, scoring inputs, query integrity
Strict-schema judge handling; escape battery survived 19/20
Scorer residency, verdict protocol, run authorization
Safety-score aggregation under judge failure
Scoring-privilege isolation in the hardened scoring library
Reward channel and verifier integrity
Verifier session and completion-channel integrity in the v1 harness
Judge-response handling in default metric paths
Built-in judge verdict handling and optimizer surface
Judges, arenas, and leaderboards
Verdict parsing and battle accounting
MT-bench extraction and arena vote processing
Judge extraction and section denominators
Injection-task scoring semantics
Aggregation weighting over published run data
Scanner and judge injection surfaces
Vote-ingestion authentication and rating integrity
Benchmark token accounting and warmup effects
Safety and guardrail tooling
Judge transcript integrity
Statuses: fix merged means a maintainer landed a fix traceable to our report; report filed means a public receipt exists and awaits action; read, unfiled findings means reproduced defects exist in our private record and appear in the ledger without a name; clean read means we read it adversarially and it held. The full defect-level record is the evidence table.