The Ledger
Every defect we have filed, every fix that landed, every break we have proven and not yet disclosed, and every target we read and found sound. Updated in weekly batches.
Three kinds of entries. Filed entries name the target and link the public receipt. Proven, unfiled entries describe the break and withhold the target; when the report is filed, the entry graduates to the named log. Sealed entries are filed privately under coordinated disclosure and unlock when the window lifts.
Landed fixes
Every row resolves to a live URL. Where the maintainer wrote or reworked the fix themselves, that is recorded too; it is the stronger outcome.
Reset-scope hole: agent-written test config survives the harness reset and grades its own pass markers
Maintainer reworked and merged the fix
Unicode normalization false negatives in detectors
Maintainer fixes #1884/#1937/#1938 merged
Scorer conflates could-not-score with attack-failed
Maintainer fix #2083 merged
Per-key validation metrics lost in aggregation
Our PR, merged by J.J. Allaire
SimpleQA grader parse brittleness alters grades
Our PR, merged
Model-graded verdict extraction trusts candidate text
Maintainer fix #4297 merged
Errored samples silently excluded from reported mean
Maintainer fix #4288 merged
Checkpoint collapse across evaluation steps
Our PR, merged
Return-value deception in judge postprocess
Maintainer fix #2565 merged
Transcript-mediated behavioral evasion class
Maintainer fix, closed completed
Chain-of-thought content unsanitized in monitor prompt
Our PR, merged
Tool-call arguments embedded unsanitized via f-string
Maintainer fix #798 merged
Model output embedded unsanitized in judge prompt
Maintainer fix #3690 merged
First-JSON extraction trusts candidate-written object
Maintainer fix #2434 merged
Filed, in the queue
Reports and patches with public receipts, awaiting maintainer action. A selection; the full set is in the evidence table.
Judge JSON parser returns the first object in the string; a planted object outranks the judge's real verdict
Grader-defect class in the reference eval suite
Verdict-extraction steering in the win-rate pipeline
Judge-output parsing defect
Grader-integrity defect in the desktop-agent benchmark
Code extraction takes first block, enabling verdict injection
Answer-extraction gaming; whitespace-brittle exact match; cache key ignores generation config
Anti-cheat canary and Dockerfile checks are integrity artifacts the evaluated policy can subvert
Verdict injection via first-JSON extraction
Verdict injection via first-JSON extraction
MixEval judge extractors take first match
Evaluation context leakage: agent can detect it is being evaluated
Proven, unfiled (134 and counting)
Breaks we have reproduced end-to-end and chosen not to file yet, or not to file at all. The target is withheld, and so is the mechanism; this page is the record of the work, not the work itself. Every entry below is PoC'd against the genuine code and traced. If you recognize your pipeline in one of these, or you want the mechanism under NDA, write to us.
The grader trusts an equality method defined by the model's own code
Code-generation evaluation, three independent harnesses·PoC'd and traced·Unfiled
Model code and its own tests run in one interpreter
Code-generation evaluation·PoC'd and traced·Unfiled
The timeout alarm is defeatable by the code it times
Code-execution evaluation·PoC'd and traced·Unfiled
Code scheduled to run at interpreter exit turns failures into recorded passes
Coding-agent benchmark·PoC'd and traced·Unfiled
Edit parsers accept non-verbatim edits without telling anyone
Coding-agent benchmark·PoC'd and traced·Unfiled
The scoring step evaluates model text as code
Math grading in a widely used harness·PoC'd and traced·Unfiled
The decoder evaluates model-supplied expressions at scoring time
Function-calling benchmark·PoC'd and traced·Unfiled
The pass/fail assertion is text the model can rewrite
Code-execution evaluation·PoC'd and traced·Unfiled
Model-written files survive the harness's own reset
Software-engineering benchmark·PoC'd and traced·Unfiled
An unanchored first-letter regex grades a refusal as correct
Reference eval suite·PoC'd and traced·Unfiled
A grader's explanation is spliced into template source
Reference eval suite·PoC'd and traced·Unfiled
The first bracketed token anywhere overrides the judge's stated verdict
Judge-comparison harness·PoC'd and traced·Unfiled
Answer-side delimiters restructure the judge prompt itself
Judge-comparison harness·PoC'd and traced·Unfiled
One unparseable verdict silently deletes a whole comparison
Judge-comparison harness·PoC'd and traced·Unfiled
The earliest match in the text steers the verdict
Win-rate pipeline·PoC'd and traced·Unfiled
Sequential placeholder substitution lets model output hijack the judge prompt
Win-rate pipeline·PoC'd and traced·Unfiled
Partially visible logprobs polarize the weighted verdict
Win-rate pipeline·PoC'd and traced·Unfiled
Plain-text speaker headers let the model re-attribute its own answers
Safety red-team framework·PoC'd and traced·Unfiled
The first digit anywhere in the response becomes the score
Vision-judge evaluation·PoC'd and traced·Unfiled
First-match verdict markers are steerable by the judged text
Broad-coverage eval suite·PoC'd and traced·Unfiled
The submission defines its own denominator
ML competition grading·PoC'd and traced·Unfiled
Judge-parse failure silently shrinks the safety denominator
Safety-certification pipeline·PoC'd and traced·Unfiled
Crashing the evaluator removes the task from the agent's own average
GUI-agent benchmark·PoC'd and traced·Unfiled
Missing subsets exit the section denominator
Reward-model benchmark·PoC'd and traced·Unfiled
Dropped predictions inflate per-length averages
Long-context benchmark·PoC'd and traced·Unfiled
Failed samples exit the pass-rate denominator
Coding-agent benchmark·PoC'd and traced·Unfiled
Unscored outputs exit the attack-success denominator
Safety scanner·PoC'd and traced·Unfiled
An error skip path records as success
Agent-safety benchmark·PoC'd and traced·Unfiled
Unparseable judge replies shrink the win-rate denominator
Win-rate pipeline·PoC'd and traced·Unfiled
The aggregation takes the maximum over per-class scores
Widely used eval harness·PoC'd and traced·Unfiled
An empty judge verdict maps to a perfect score on nine safety metrics
Safety-evaluation library·PoC'd and traced·Unfiled
A perfect score is mathematically unreachable on multi-answer rows
Legal-reasoning benchmark·PoC'd and traced·Unfiled
A zero-valued gold answer is mathematically unwinnable
Legal-reasoning benchmark·PoC'd and traced·Unfiled
A negated statement scores as correct under containment matching
Legal-reasoning benchmark·PoC'd and traced·Unfiled
A fixed two-word hedge scores a perfect 100
Long-context benchmark·PoC'd and traced·Unfiled
Gold strings match inside unrelated words
Long-context benchmark·PoC'd and traced·Unfiled
Percentage leniency equates answers a hundred times apart
Math grading·PoC'd and traced·Unfiled
Format-equivalent answers are scored incorrect
Math grading·PoC'd and traced·Unfiled
The grader evaluates the echoed problem instead of the answer
Math verification library·PoC'd and traced·Unfiled
Percent-formatted golds can never be matched
Math verification library·PoC'd and traced·Unfiled
Fuzzy containment passes near-misses as exact
Reference eval suite·PoC'd and traced·Unfiled
The response cache key omits model identity
Widely used eval harness·PoC'd and traced·Unfiled
A first-sentinel extractor trusts the wrong answer marker
Math grading·PoC'd and traced·Unfiled
Duplicate items give one question outsized metric leverage
A lab's public eval data·PoC'd and traced·Unfiled
Forged conversation turns sit inside eval questions
A lab's public eval data·PoC'd and traced·Unfiled
An always-truthy injection check awards full objective credit
Agent-safety benchmark, two task versions·PoC'd and traced·Unfiled
Zero-work agents earn a perfect reward on shipped tasks
Conversational-agent benchmark·PoC'd and traced·Unfiled
The user simulator accepts injected termination
Conversational-agent benchmark·PoC'd and traced·Unfiled
The grader replays gold actions through the live environment
Conversational-agent benchmark·PoC'd and traced·Unfiled
An allowlisted calculator can hang the whole evaluation
Conversational-agent benchmark·PoC'd and traced·Unfiled
A timeout appends a passing grade on every infeasible task
GUI-agent benchmark·PoC'd and traced·Unfiled
Gold tests import agent-written modules on the grader host
GUI-agent benchmark·PoC'd and traced·Unfiled
The entrant can reset the competition's retry guard
Prize competition·PoC'd and traced·Unfiled
Prefix-colliding identifiers shrink the scorecard denominator
Prize competition·PoC'd and traced·Unfiled
One malformed tool call crashes the entire benchmark run
Agent-safety benchmark·PoC'd and traced·Unfiled
Task-suite includes read arbitrary local files
Agent-safety benchmark·PoC'd and traced·Unfiled
A post-start permission step hands the agent ownership of the official scorer
Research-lab agent harness·PoC'd and traced·Unfiled
Score verdicts ride on output behind a public sentinel
Research-lab agent harness·PoC'd and traced·Unfiled
Hidden scores leak their reasoning to the agent
Research-lab agent harness·PoC'd and traced·Unfiled
One agent's token can score and poison other runs
Research-lab agent harness·PoC'd and traced·Unfiled
The reward file is forged through a PATH shim
Terminal-task benchmark·PoC'd and traced·Unfiled
CI executes submission-controlled configuration on the runner
Benchmark repository CI·PoC'd and traced·Unfiled
An LLM's review output is substituted into a shell step
Benchmark repository CI·PoC'd and traced·Unfiled
A canary check silently skips tasks with spaces in their names
Benchmark repository CI·PoC'd and traced·Unfiled
Self-reported score files land with no verification gate
Public leaderboard pipeline·PoC'd and traced·Unfiled
Payloads placed only in losing answers cause the losses to be discarded
Arena vote processing·PoC'd and traced·Unfiled
A vote filter purges honest voters after one dissent
Arena vote processing·PoC'd and traced·Unfiled
The repo ships committed live session credentials
Web-agent benchmark·PoC'd and traced·Unfiled
Annotators execute model-generated SQL unsandboxed
Leaderboard eval suite·PoC'd and traced·Unfiled
Model output writes arbitrary files that a build step then executes
Leaderboard eval suite·PoC'd and traced·Unfiled
Judged text is evaluated with full builtins
Broad-coverage eval suite·PoC'd and traced·Unfiled
An external content API's empty response is recorded as all-safe
Safety pipeline·PoC'd and traced·Unfiled
Managed judge templates splice candidate output inline
Eval observability platform·PoC'd and traced·Unfiled
Prototype-chain keys splice into judge context
Eval observability platform·PoC'd and traced·Unfiled
Sorting parameters inject into analytic queries
Eval observability platform·PoC'd and traced·Unfiled
Task output shadows ground truth in scoring inputs
Eval observability platform·PoC'd and traced·Unfiled
Terminal hyperlinks can carry script URLs into logs
Eval observability platform·PoC'd and traced·Unfiled
A rubric judge parses free text with no structured enforcement
Interpretability tooling·PoC'd and traced·Unfiled
A login shell sources a user-writable profile before the official scorer runs
Protected-scoring library·PoC'd and traced·Unfiled
Nothing checks the official scorer's integrity at scoring time
Protected-scoring library·PoC'd and traced·Unfiled
The privileged scorer imports the agent's own module to grade it
Research-lab benchmark suite·PoC'd and traced·Unfiled
Agent data artifacts are deserialized inside the privileged scoring process
Research-lab benchmark suite·PoC'd and traced·Unfiled
A dropped package shadows the scoring library inside its own process
Research-lab benchmark suite·PoC'd and traced·Unfiled
The scoring equation leaks its ground truth to the agent
Research-lab benchmark suite·PoC'd and traced·Unfiled
A clock the scorer trusts is settable by the agent
Research-lab benchmark suite·PoC'd and traced·Unfiled
The grader imports the agent's kernel and lets it define equality
Kernel-optimization evaluation·PoC'd and traced·Unfiled
The reward-hacking scanner is injectable by the transcript it scans
Agent-monitoring tooling·PoC'd and traced·Unfiled
A judge faithfully quoting a poisoned submission records the poisoned decision
Agent-monitoring tooling·PoC'd and traced·Unfiled
Secret redaction skips the metadata fields
Agent-monitoring tooling·PoC'd and traced·Unfiled
Task-equal weighting silently weights by run count; published ranks flip
Evaluation analytics toolkit·PoC'd and traced·Unfiled
The scorer builds a root shell from agent submission bytes
Research-lab task family·PoC'd and traced·Unfiled
The last printed line decides the recorded grade
Research-lab task family·PoC'd and traced·Unfiled
A resource sampler misses processes that live briefly
Research-lab task family·PoC'd and traced·Unfiled
An unauthenticated in-container oracle hands over the secret answer
Research-lab agent harness·PoC'd and traced·Unfiled
Agent text lands as the innermost turn of the user simulator's prompt
Conversational-agent benchmark, successor version·PoC'd and traced·Unfiled
The successor benchmark inherits free-reward tasks and adds an injection-driven termination chain
Conversational-agent benchmark, successor version·PoC'd and traced·Unfiled
Golden actions are credited to user-side calls under dual control
Conversational-agent benchmark, successor version·PoC'd and traced·Unfiled
A non-preemptive timeout lets a six-second step survive a two-second limit
Conversational-agent benchmark, successor version·PoC'd and traced·Unfiled
Unparseable judge rows silently exit the headline error metric
Enterprise-agent evaluation suite·PoC'd and traced·Unfiled
A schema-valid empty judge reply collapses the error metric to zero
Enterprise-agent evaluation suite·PoC'd and traced·Unfiled
A pre-created terminal session captures the verifier's keystrokes and answers them
Terminal-task benchmark·PoC'd and traced·Unfiled
The verifier prefers a pre-seeded reward file over the failing test verdict
Terminal-task benchmark, next generation·PoC'd and traced·Unfiled
The verifier's own reward write imports an agent-planted module
Benchmark platform·PoC'd and traced·Unfiled
The verifier installs the agent's dependency list as root before grading
Terminal-task benchmark, science variant·PoC'd and traced·Unfiled
A shell tool sources the working directory's env file before parsing arguments
Benchmark platform tooling·PoC'd and traced·Unfiled
Stuffing trials on one easy task moves the headline metric
Benchmark indexing pipeline·PoC'd and traced·Unfiled
Metric text is misread as a test verdict
Terminal-task benchmark, next generation·PoC'd and traced·Unfiled
The MCQ grader reads a line-start letter as the answer
Frontier open-weight model's vendored harness·PoC'd and traced·Unfiled
An early answer marker overrides the true final answer
Frontier open-weight model's vendored harness·PoC'd and traced·Unfiled
A literal placeholder captures the rubric into the answer region
Frontier open-weight model's vendored harness·PoC'd and traced·Unfiled
The numeric grader evaluates the model's extracted answer as code
Open-weight math model's grading suite·PoC'd and traced·Unfiled
A symbolic-equivalence check executes attribute chains in model text
Open-weight math model's grading suite·PoC'd and traced·Unfiled
Equivalence defaults pass hundredfold misses and prefix-dropped answers
Open-weight math model's grading suite·PoC'd and traced·Unfiled
The code-executor sandbox is two substring checks
Open-weight model's program-of-thought grading·PoC'd and traced·Unfiled
A formal-proof grader returns True unconditionally
Formal-math evaluation·PoC'd and traced·Unfiled
The access-control oracle reimplements the policy with a divergence
Multitenant evaluation suite·PoC'd and traced·Unfiled
The task recomputes ground truth from the artifact the agent just edited
Enterprise workflow benchmark·PoC'd and traced·Unfiled
The URL gate ignores the query string the answer depends on
Enterprise workflow benchmark·PoC'd and traced·Unfiled
Shared-instance admin credentials ship behind a repo-hardcoded key
Enterprise workflow benchmark·PoC'd and traced·Unfiled
Agent-overwritable result artifacts execute on the analyst's machine
Enterprise workflow benchmark·PoC'd and traced·Unfiled
Verdict injection into the curation judge poisons the training set
Function-calling data pipeline·PoC'd and traced·Unfiled
The score-threshold gate is plumbed but never read
Function-calling data pipeline·PoC'd and traced·Unfiled
Throughput is counted from requested tokens, not generated ones
Serving benchmark used industry-wide·PoC'd and traced·Unfiled
The warmup primes the first measured request byte-identically
Serving benchmark used industry-wide·PoC'd and traced·Unfiled
Unauthenticated vote ingestion manufactures rating points
Public model arena·PoC'd and traced·Unfiled
Password reset binds the wrong argument
Public model arena·PoC'd and traced·Unfiled
Malformed votes silently corrupt the rating fit
Arena ranking pipeline·PoC'd and traced·Unfiled
The default metric path evaluates the raw judge response as code
Enterprise RAG evaluation library·PoC'd and traced·Unfiled
Off-enum judge verdicts fail open and exit the mean
ML platform's built-in judges·PoC'd and traced·Unfiled
Model output drives a regex over the full trace
ML platform's built-in judges·PoC'd and traced·Unfiled
The optimizer distills raw outputs into deployed judge guidelines
ML platform's optimization tooling·PoC'd and traced·Unfiled
An LLM's reason string lands in a privileged shell step on fork PRs
Evaluation library CI·PoC'd and traced·Unfiled
An issue title lands in a shell step, inert only behind a disabled job
AI developer-tool CI·PoC'd and traced·Unfiled
Four more pipelines sit one config change from the same primitive
Benchmark and agent-framework CI·PoC'd and traced·Unfiled
Sealed
Filed privately with the affected organizations under coordinated disclosure. These unlock when their windows lift.
Clean reads
Targets we read adversarially and found sound, or found hardened far past the norm. Publishing these is what makes the findings above worth trusting.
Judge-output handling is strict-schema end to end; a 20-part escape battery produced one medium finding. The closest thing to a control case we have read.
All 2,400 human labels join cleanly to the dataset; numeric re-verification found the labels defensible; the paper's headline failure rate reproduces exactly from the shipped data.
Template-rendering surface correctly delegated and sandboxed. No finding.
Chat-template sandboxing is textbook. No finding.
Standing studies
Daily hashes of the public datasets, leaderboards, and grading code the industry's scores rest on. Append-only, running since August 2026.
The ETB taxonomyThe ten mechanisms behind 76 of the first 99 defects, with counts and fixes.
The paperThe evaluator trust boundary, the instrument-residency hierarchy, and the live judge studies. Peer-archived, DOI.
Counts on this page are conservative and traceable; where a defect produced both an issue and a pull request it is counted once. Corrections welcome: john@authensor.com.