Public receipts for every one: reproductions, file and line, upstream issues. A 24th is under coordinated disclosure until November. The one clean repository does the one thing the other 24 don't: it validates the judge's structured output instead of parsing free text.
The 25 most-starred open AI evaluation repositories, snapshot 2026-08-15. For each, we read the grading path end to end: dataset residency, answer extraction, judge parsing, aggregation, reporting. No surface scanners. The path that turns a model's output into a number, read line by line.
A defect counts only when it is proven, not asserted. Every red square has a standalone reproduction that downloads the genuine upstream file at a pinned commit, verifies the sha256, and runs the real parser against a crafted candidate output. If the score moves, the square goes red. If it doesn't, it doesn't.
The defects are not exotic. The first JSON object in judge free text wins. The word "Yes" anywhere in judge text means correct. The last two numbers in the reply are the score. Ground truth read from inside the machine being graded.
Statuses change. When a fix lands upstream, the ledger records it and the next snapshot reflects it. The point of the census is not that the industry is careless; it is that the failure is one architectural decision, repeated 24 times.
More on the ledger: landed fixes, proven-unfiled shapes with targets withheld, and the sealed count.
The clean repository validates the judge's structured output: verdicts arrive as typed fields, labels are checked against the allowed set, and malformed output fails loud instead of being parsed into a passing grade. That is the whole fix. It is a one-file architectural decision, and 24 of 25 made the other one.
The class is not inevitable. That is what makes it worth fixing, and what makes it worth auditing. Read the taxonomy for the 75 ways the boundary fails, or the ledger for what happened when we told the maintainers.