The audited surface

The Map

The instruments we have read, what we read them for, and where each one stands. A finding is named only where a public receipt exists; everything else carries a status label, not a mechanism.

How we read

Line by line

We read the scoring path, not the README. The grader, the parser, the aggregation, the cache, the reset.

Proven or dropped

Every candidate defect must reproduce against the genuine code path. Hypotheses that fail to execute are recorded as refuted and never reported.

Deduped against the record

Every finding is checked against the project's own tracker and the public literature before it is counted. Known issues are cited, not re-claimed.

Private first

Reports go to the maintainer before anywhere else, with the reproduction and a fix. Publication follows on the vendor's clock.

Benchmarks

princeton-nlp/SWE-benchfix merged ↗

Reset-scope and grading-path integrity

openai/mle-benchread, unfiled findings

Submission and denominator integrity

sierra-research/tau-benchread, unfiled findings

Reward pipeline and task-suite integrity

web-arena-x/webarenaread, unfiled findings

Evaluator crash paths and aggregation

xlang-ai/OSWorldreport filed ↗

Grader-host integrity and timeout semantics

aymeric-roucher/GAIAprivate disclosure pending

Scorer quirks and repository hygiene

arcprize/ARC-AGIread, unfiled findings

Competition guard and scorecard integrity

HazyResearch/legalbenchread, unfiled findings

Metric correctness under multi-answer scoring

patronus-ai/financebenchclean read

Labels re-verified; headline reproduces exactly

LiveBench/LiveBenchread, unfiled findings

Coding grader residency and judgment integrity

NVIDIA/RULERread, unfiled findings

String-matching and denominator semantics

openai/simple-evalsreport filed ↗

Grader parse and template integrity

METR/RE-Benchread, unfiled findings

Protected-scoring integrity across all seven task families

METR/public-tasksread, unfiled findings

Task-family scorer integrity census

sierra-research/tau2-benchread, unfiled findings

Successor-suite reward pipeline and user-simulator integrity

sierra-research/mu-benchread, unfiled findings

Headline metric aggregation under judge failure

openai/gpt-ossread, unfiled findings

Vendored harness grader and rubric-template integrity

deepseek-ai/DeepSeek-Mathread, unfiled findings

Grader execution and equivalence semantics

deepseek-ai/DeepSeek-Coderread, unfiled findings

Pass@k and executor grading integrity

ServiceNow/WorkArenaread, unfiled findings

Retrieval-task ground-truth integrity

ServiceNow/BrowserGymread, unfiled findings

Result artifact integrity

salesforce/xLAMread, unfiled findings

Curation-judge and data-pipeline integrity

ogx-ai/ogxread, unfiled findings

Multitenant evaluation suite integrity

Harnesses and scoring libraries

EleutherAI/lm-evaluation-harnessPRs open ↗

Answer extraction, caching, math grading

UKGovernmentBEIS/inspect_aifixes merged ↗

Verdict extraction, error handling, judge prompts

UKGovernmentBEIS/inspect_evalsfix merged ↗

Per-task grader integrity

stanford-crfm/helmread, unfiled findings

Annotator execution and score extraction

huggingface/lightevalPR open ↗

Judge extractors and math scoring

open-compass/opencompassfix merged ↗

Judge postprocess and evaluated-text execution

confident-ai/deepevalPR open ↗

Verdict parser integrity

promptfoo/promptfooPR open ↗

Judge JSON extraction

comet-ml/opikreport filed, maintainer fix in review ↗

Judge parsing, scoring inputs, query integrity

Arize-ai/phoenixclean read (near)

Strict-schema judge handling; escape battery survived 19/20

METR/vivariaread, unfiled findings

Scorer residency, verdict protocol, run authorization

mlcommons/modelgaugeread, unfiled findings

Safety-score aggregation under judge failure

METR/task-protected-scoringread, unfiled findings

Scoring-privilege isolation in the hardened scoring library

harbor-framework/harborread, unfiled findings

Reward channel and verifier integrity

Mercor-io/terminal-benchread, unfiled findings

Verifier session and completion-channel integrity in the v1 harness

truera/trulensread, unfiled findings

Judge-response handling in default metric paths

mlflow/mlflow (genai)read, unfiled findings

Built-in judge verdict handling and optimizer surface

Judges, arenas, and leaderboards

lmarena/arena-hard-autoread, unfiled findings

Verdict parsing and battle accounting

lm-sys/FastChatread, unfiled findings

MT-bench extraction and arena vote processing

allenai/reward-benchread, unfiled findings

Judge extraction and section denominators

tatsu-lab/alpaca_evalreport filed ↗

Win-rate construction and parser steering

Mercor-io/terminal-bench-3reports filed ↗

Leaderboard verification and anti-cheat integrity

ethz-spylab/agentdojoread, unfiled findings

Injection-task scoring semantics

METR/eval-analysis-publicread, unfiled findings

Aggregation weighting over published run data

METR/metr-inspect-agentsread, unfiled findings

Scanner and judge injection surfaces

lmarena/copilot-arenaread, unfiled findings

Vote-ingestion authentication and rating integrity

sgl-project/sglang (benchmark)read, unfiled findings

Benchmark token accounting and warmup effects

Safety and guardrail tooling

NVIDIA/garakfixes merged ↗

Detector normalization and ASR accounting

microsoft/PyRITfix merged ↗

Scorer semantics for could-not-score

UKGovernmentBEIS/control-arenafixes merged ↗

Monitor prompt construction

meridianlabs-ai/inspect_scoutfix merged ↗

Validation-metric aggregation

Giskard-AI/giskardread, unfiled findings

Judge transcript integrity

meta-llama/PurpleLlamaPR open ↗

Scanner fail-open semantics

Statuses: fix merged means a maintainer landed a fix traceable to our report; report filed means a public receipt exists and awaits action; read, unfiled findings means reproduced defects exist in our private record and appear in the ledger without a name; clean read means we read it adversarially and it held. The full defect-level record is the evidence table.