Every defect we have filed, every fix that landed, every break we have proven and not yet disclosed, and every target we read and found sound. Updated in weekly batches.
Every row resolves to a live URL. Where the maintainer wrote or reworked the fix themselves, that is recorded too; it is the stronger outcome.
Reset-scope hole: agent-written test config survives the harness reset and grades its own pass markers
Maintainer reworked and merged the fix
Unicode normalization false negatives in detectors
Maintainer fixes #1884/#1937/#1938 merged
Scorer conflates could-not-score with attack-failed
Maintainer fix #2083 merged
SSN over-blocking false positive
Our PR, merged
NIF/NIE checksum validation gap
Our PR, merged
Punycode email recognition gap
Our PR, merged
IBAN format coverage gap
Our PR, merged
Per-key validation metrics lost in aggregation
Our PR, merged by J.J. Allaire
SimpleQA grader parse brittleness alters grades
Our PR, merged
DNS-rebinding SSRF in model serving
Maintainer fix #24258 merged
Model-graded verdict extraction trusts candidate text
Maintainer fix #4297 merged
Errored samples silently excluded from reported mean
Maintainer fix #4288 merged
Checkpoint collapse across evaluation steps
Our PR, merged
Return-value deception in judge postprocess
Maintainer fix #2565 merged
Legacy ingestion path bypasses masking
Maintainer fix merged
Transcript-mediated behavioral evasion class
Maintainer fix, closed completed
Chain-of-thought content unsanitized in monitor prompt
Our PR, merged
Tool-call arguments embedded unsanitized via f-string
Maintainer fix #798 merged
Model output embedded unsanitized in judge prompt
Maintainer fix #3690 merged
Command injection in pattern handling
Maintainer merged
First-JSON extraction trusts candidate-written object
Maintainer fix #2434 merged
Security defect in workflow loading
Maintainer fix #14774 merged
Reports and patches with public receipts, awaiting maintainer action. A selection; the full set is in the evidence table.
Judge JSON parser returns the first object in the string; a planted object outranks the judge's real verdict
Grader-defect class in the reference eval suite
Verdict-extraction steering in the win-rate pipeline
Judge-output parsing defect
Grader-integrity defect in the desktop-agent benchmark
Code extraction takes first block, enabling verdict injection
Answer-extraction gaming; whitespace-brittle exact match; cache key ignores generation config
Anti-cheat canary and Dockerfile checks are integrity artifacts the evaluated policy can subvert
Verdict injection via first-JSON extraction
Verdict injection via first-JSON extraction
MixEval judge extractors take first match
Evaluation context leakage: agent can detect it is being evaluated
Breaks we have reproduced end-to-end and chosen not to file yet, or not to file at all. The target is withheld, and so is the mechanism; this page is the record of the work, not the work itself. Every entry below is PoC'd against the genuine code and traced. If you recognize your pipeline in one of these, or you want the mechanism under NDA, write to us.
The grader trusts an equality method defined by the model's own code
Code-generation evaluation, three independent harnessesModel code and its own tests run in one interpreter
Code-generation evaluationThe timeout alarm is defeatable by the code it times
Code-execution evaluationCode scheduled to run at interpreter exit turns failures into recorded passes
Coding-agent benchmarkEdit parsers accept non-verbatim edits without telling anyone
Coding-agent benchmarkThe scoring step evaluates model text as code
Math grading in a widely used harnessThe decoder evaluates model-supplied expressions at scoring time
Function-calling benchmarkThe pass/fail assertion is text the model can rewrite
Code-execution evaluationModel-written files survive the harness's own reset
Software-engineering benchmarkAn unanchored first-letter regex grades a refusal as correct
Reference eval suiteA grader's explanation is spliced into template source
Reference eval suiteThe first bracketed token anywhere overrides the judge's stated verdict
Judge-comparison harnessAnswer-side delimiters restructure the judge prompt itself
Judge-comparison harnessOne unparseable verdict silently deletes a whole comparison
Judge-comparison harnessThe earliest match in the text steers the verdict
Win-rate pipelineSequential placeholder substitution lets model output hijack the judge prompt
Win-rate pipelinePartially visible logprobs polarize the weighted verdict
Win-rate pipelinePlain-text speaker headers let the model re-attribute its own answers
Safety red-team frameworkThe first digit anywhere in the response becomes the score
Vision-judge evaluationFirst-match verdict markers are steerable by the judged text
Broad-coverage eval suiteThe submission defines its own denominator
ML competition gradingJudge-parse failure silently shrinks the safety denominator
Safety-certification pipelineCrashing the evaluator removes the task from the agent's own average
GUI-agent benchmarkMissing subsets exit the section denominator
Reward-model benchmarkDropped predictions inflate per-length averages
Long-context benchmarkFailed samples exit the pass-rate denominator
Coding-agent benchmarkUnscored outputs exit the attack-success denominator
Safety scannerAn error skip path records as success
Agent-safety benchmarkUnparseable judge replies shrink the win-rate denominator
Win-rate pipelineThe aggregation takes the maximum over per-class scores
Widely used eval harnessAn empty judge verdict maps to a perfect score on nine safety metrics
Safety-evaluation libraryA perfect score is mathematically unreachable on multi-answer rows
Legal-reasoning benchmarkA zero-valued gold answer is mathematically unwinnable
Legal-reasoning benchmarkA negated statement scores as correct under containment matching
Legal-reasoning benchmarkA fixed two-word hedge scores a perfect 100
Long-context benchmarkGold strings match inside unrelated words
Long-context benchmarkPercentage leniency equates answers a hundred times apart
Math gradingFormat-equivalent answers are scored incorrect
Math gradingThe grader evaluates the echoed problem instead of the answer
Math verification libraryPercent-formatted golds can never be matched
Math verification libraryFuzzy containment passes near-misses as exact
Reference eval suiteThe response cache key omits model identity
Widely used eval harnessA first-sentinel extractor trusts the wrong answer marker
Math gradingDuplicate items give one question outsized metric leverage
A lab's public eval dataForged conversation turns sit inside eval questions
A lab's public eval dataAn always-truthy injection check awards full objective credit
Agent-safety benchmark, two task versionsZero-work agents earn a perfect reward on shipped tasks
Conversational-agent benchmarkThe user simulator accepts injected termination
Conversational-agent benchmarkThe grader replays gold actions through the live environment
Conversational-agent benchmarkAn allowlisted calculator can hang the whole evaluation
Conversational-agent benchmarkA timeout appends a passing grade on every infeasible task
GUI-agent benchmarkGold tests import agent-written modules on the grader host
GUI-agent benchmarkThe entrant can reset the competition's retry guard
Prize competitionPrefix-colliding identifiers shrink the scorecard denominator
Prize competitionOne malformed tool call crashes the entire benchmark run
Agent-safety benchmarkTask-suite includes read arbitrary local files
Agent-safety benchmarkA post-start permission step hands the agent ownership of the official scorer
Research-lab agent harnessScore verdicts ride on output behind a public sentinel
Research-lab agent harnessHidden scores leak their reasoning to the agent
Research-lab agent harnessOne agent's token can score and poison other runs
Research-lab agent harnessThe reward file is forged through a PATH shim
Terminal-task benchmarkCI executes submission-controlled configuration on the runner
Benchmark repository CIAn LLM's review output is substituted into a shell step
Benchmark repository CIA canary check silently skips tasks with spaces in their names
Benchmark repository CISelf-reported score files land with no verification gate
Public leaderboard pipelinePayloads placed only in losing answers cause the losses to be discarded
Arena vote processingA vote filter purges honest voters after one dissent
Arena vote processingThe repo ships committed live session credentials
Web-agent benchmarkAnnotators execute model-generated SQL unsandboxed
Leaderboard eval suiteModel output writes arbitrary files that a build step then executes
Leaderboard eval suiteJudged text is evaluated with full builtins
Broad-coverage eval suiteAn external content API's empty response is recorded as all-safe
Safety pipelineManaged judge templates splice candidate output inline
Eval observability platformPrototype-chain keys splice into judge context
Eval observability platformSorting parameters inject into analytic queries
Eval observability platformTask output shadows ground truth in scoring inputs
Eval observability platformTerminal hyperlinks can carry script URLs into logs
Eval observability platformA rubric judge parses free text with no structured enforcement
Interpretability toolingA login shell sources a user-writable profile before the official scorer runs
Protected-scoring libraryNothing checks the official scorer's integrity at scoring time
Protected-scoring libraryThe privileged scorer imports the agent's own module to grade it
Research-lab benchmark suiteAgent data artifacts are deserialized inside the privileged scoring process
Research-lab benchmark suiteA dropped package shadows the scoring library inside its own process
Research-lab benchmark suiteThe scoring equation leaks its ground truth to the agent
Research-lab benchmark suiteA clock the scorer trusts is settable by the agent
Research-lab benchmark suiteThe grader imports the agent's kernel and lets it define equality
Kernel-optimization evaluationThe reward-hacking scanner is injectable by the transcript it scans
Agent-monitoring toolingA judge faithfully quoting a poisoned submission records the poisoned decision
Agent-monitoring toolingSecret redaction skips the metadata fields
Agent-monitoring toolingTask-equal weighting silently weights by run count; published ranks flip
Evaluation analytics toolkitThe scorer builds a root shell from agent submission bytes
Research-lab task familyThe last printed line decides the recorded grade
Research-lab task familyA resource sampler misses processes that live briefly
Research-lab task familyAn unauthenticated in-container oracle hands over the secret answer
Research-lab agent harnessAgent text lands as the innermost turn of the user simulator's prompt
Conversational-agent benchmark, successor versionThe successor benchmark inherits free-reward tasks and adds an injection-driven termination chain
Conversational-agent benchmark, successor versionGolden actions are credited to user-side calls under dual control
Conversational-agent benchmark, successor versionA non-preemptive timeout lets a six-second step survive a two-second limit
Conversational-agent benchmark, successor versionUnparseable judge rows silently exit the headline error metric
Enterprise-agent evaluation suiteA schema-valid empty judge reply collapses the error metric to zero
Enterprise-agent evaluation suiteA pre-created terminal session captures the verifier's keystrokes and answers them
Terminal-task benchmarkThe verifier prefers a pre-seeded reward file over the failing test verdict
Terminal-task benchmark, next generationThe verifier's own reward write imports an agent-planted module
Benchmark platformThe verifier installs the agent's dependency list as root before grading
Terminal-task benchmark, science variantA shell tool sources the working directory's env file before parsing arguments
Benchmark platform toolingStuffing trials on one easy task moves the headline metric
Benchmark indexing pipelineMetric text is misread as a test verdict
Terminal-task benchmark, next generationThe MCQ grader reads a line-start letter as the answer
Frontier open-weight model's vendored harnessAn early answer marker overrides the true final answer
Frontier open-weight model's vendored harnessA literal placeholder captures the rubric into the answer region
Frontier open-weight model's vendored harnessThe numeric grader evaluates the model's extracted answer as code
Open-weight math model's grading suiteA symbolic-equivalence check executes attribute chains in model text
Open-weight math model's grading suiteEquivalence defaults pass hundredfold misses and prefix-dropped answers
Open-weight math model's grading suiteThe code-executor sandbox is two substring checks
Open-weight model's program-of-thought gradingA formal-proof grader returns True unconditionally
Formal-math evaluationThe access-control oracle reimplements the policy with a divergence
Multitenant evaluation suiteThe task recomputes ground truth from the artifact the agent just edited
Enterprise workflow benchmarkThe URL gate ignores the query string the answer depends on
Enterprise workflow benchmarkShared-instance admin credentials ship behind a repo-hardcoded key
Enterprise workflow benchmarkAgent-overwritable result artifacts execute on the analyst's machine
Enterprise workflow benchmarkVerdict injection into the curation judge poisons the training set
Function-calling data pipelineThe score-threshold gate is plumbed but never read
Function-calling data pipelineThroughput is counted from requested tokens, not generated ones
Serving benchmark used industry-wideThe warmup primes the first measured request byte-identically
Serving benchmark used industry-wideUnauthenticated vote ingestion manufactures rating points
Public model arenaPassword reset binds the wrong argument
Public model arenaMalformed votes silently corrupt the rating fit
Arena ranking pipelineThe default metric path evaluates the raw judge response as code
Enterprise RAG evaluation libraryOff-enum judge verdicts fail open and exit the mean
ML platform's built-in judgesModel output drives a regex over the full trace
ML platform's built-in judgesThe optimizer distills raw outputs into deployed judge guidelines
ML platform's optimization toolingAn LLM's reason string lands in a privileged shell step on fork PRs
Evaluation library CIAn issue title lands in a shell step, inert only behind a disabled job
AI developer-tool CIFour more pipelines sit one config change from the same primitive
Benchmark and agent-framework CIAll PoC'd and traced · all unfiled
Filed privately with the affected organizations. These unlock when their windows lift. The volume is visible; the contents are not.
Targets we read adversarially and found sound, or found hardened far past the norm. Publishing these is what makes the findings above worth trusting.
Judge-output handling is strict-schema end to end; a 20-part escape battery produced one medium finding. The closest thing to a control case we have read.
All 2,400 human labels join cleanly to the dataset; numeric re-verification found the labels defensible; the paper's headline failure rate reproduces exactly from the shipped data.
Template-rendering surface correctly delegated and sandboxed. No finding.
Chat-template sandboxing is textbook. No finding.
Daily hashes of the public datasets, leaderboards, and grading code the industry's scores rest on. Append-only, running since August 2026.
The ETB taxonomyThe ten mechanisms behind 76 of the first 99 defects, with counts and fixes.
The paperThe evaluator trust boundary, the instrument-residency hierarchy, and the live judge studies. Peer-archived, DOI.
Counts on this page are conservative and traceable; where a defect produced both an issue and a pull request it is counted once. Corrections welcome: john@authensor.com.
No charge, no strings. It is how this ledger grew.
You send one public claim. A benchmark score, a safety rate, a leaderboard position.
We read the instrument, the same way every row above was produced.
One page back: what holds, what is unverifiable, what is broken. Each with a reproduction.