23 of the 25 most-starred evaluation repositories have a trust-boundary defect.

Public receipts for every one: reproductions, file and line, upstream issues. A 24th is under coordinated disclosure until November. The one clean repository does the one thing the other 24 don't: it validates the judge's structured output instead of parsing free text.

FastChat
promptfoo
opik
openai/evals
deepeval
ragas
lm-eval-harness
phoenix
garak
opencompass
SWE-bench
simple-evals
verifiers
PurpleLlama
PyRIT
AgentBench
trulens
human-eval
evalscope
OSWorld
helm
inspect_ai
lighteval
uptrain
alpaca_eval
publicly verifiable defect (23) sealed, coordinated disclosure (1) audited clean (1)

How a square goes red.

The 25 most-starred open AI evaluation repositories, snapshot 2026-08-15. For each, we read the grading path end to end: dataset residency, answer extraction, judge parsing, aggregation, reporting. No surface scanners. The path that turns a model's output into a number, read line by line.

A defect counts only when it is proven, not asserted. Every red square has a standalone reproduction that downloads the genuine upstream file at a pinned commit, verifies the sha256, and runs the real parser against a crafted candidate output. If the score moves, the square goes red. If it doesn't, it doesn't.

The defects are not exotic. The first JSON object in judge free text wins. The word "Yes" anywhere in judge text means correct. The last two numbers in the reply are the score. Ground truth read from inside the machine being graded.

Statuses change. When a fix lands upstream, the ledger records it and the next snapshot reflects it. The point of the census is not that the industry is careless; it is that the failure is one architectural decision, repeated 24 times.

One repo made the other decision.

The clean repository validates the judge's structured output: verdicts arrive as typed fields, labels are checked against the allowed set, and malformed output fails loud instead of being parsed into a passing grade. That is the whole fix. It is a one-file architectural decision, and 24 of 25 made the other one.

The class is not inevitable. That is what makes it worth fixing, and what makes it worth auditing. Read the taxonomy for the 75 ways the boundary fails, or the ledger for what happened when we told the maintainers.