The Ledger.

Every defect we have filed, every fix that landed, every break we have proven and not yet disclosed, and every target we read and found sound. Updated in weekly batches.

Entry classes
FiledTarget named, public receipt linked. The receipt is the proof.
Proven, unfiledThe break is described and the target withheld. When the report is filed, the entry graduates to the named log.
SealedFiled privately under coordinated disclosure. Unlocks when the window lifts.
The record, to date
Read100+Repositories read line by line, across 45+ organizations
Filed200+Defect reports: pull requests and issues, every one with a public receipt
Landed30+Upstream receipts: merged, or fixed by maintainers after we filed
Sealed4Filed privately, awaiting disclosure windows
01Landed

Fixes that landed upstream.

Every row resolves to a live URL. Where the maintainer wrote or reworked the fix themselves, that is recorded too; it is the stronger outcome.

2026-08-13
princeton-nlp/SWE-bench

Reset-scope hole: agent-written test config survives the harness reset and grades its own pass markers

Maintainer reworked and merged the fix

#620
2026-07-24
NVIDIA/garak

Unicode normalization false negatives in detectors

Maintainer fixes #1884/#1937/#1938 merged

#1867
2026-07-13
microsoft/PyRIT

Scorer conflates could-not-score with attack-failed

Maintainer fix #2083 merged

#2044
2026-07-13
data-privacy-stack/presidio

SSN over-blocking false positive

Our PR, merged

#2074
2026-07-01
data-privacy-stack/presidio

NIF/NIE checksum validation gap

Our PR, merged

#2076
2026-07-01
data-privacy-stack/presidio

Punycode email recognition gap

Our PR, merged

#2077
2026-07-01
data-privacy-stack/presidio

IBAN format coverage gap

Our PR, merged

#2078
2026-06-26
meridianlabs-ai/inspect_scout

Per-key validation metrics lost in aggregation

Our PR, merged by J.J. Allaire

#473
2026-06-24
UKGovernmentBEIS/inspect_evals

SimpleQA grader parse brittleness alters grades

Our PR, merged

#1812
2026-06-24
mlflow/mlflow

DNS-rebinding SSRF in model serving

Maintainer fix #24258 merged

#24179
2026-06-18
UKGovernmentBEIS/inspect_ai

Model-graded verdict extraction trusts candidate text

Maintainer fix #4297 merged

#4283
2026-06-18
UKGovernmentBEIS/inspect_ai

Errored samples silently excluded from reported mean

Maintainer fix #4288 merged

#4286
2026-06-18
UKGovernmentBEIS/inspect_cyber

Checkpoint collapse across evaluation steps

Our PR, merged

#114
2026-06-18
open-compass/opencompass

Return-value deception in judge postprocess

Maintainer fix #2565 merged

#2535
2026-06-18
langfuse/langfuse

Legacy ingestion path bypasses masking

Maintainer fix merged

#14576
2026-06-18
UKGovernmentBEIS/sandbox_escape_bench

Transcript-mediated behavioral evasion class

Maintainer fix, closed completed

#18
2026-04-16
UKGovernmentBEIS/control-arena

Chain-of-thought content unsanitized in monitor prompt

Our PR, merged

#810
2026-03-31
UKGovernmentBEIS/control-arena

Tool-call arguments embedded unsanitized via f-string

Maintainer fix #798 merged

#791
2026-03-31
UKGovernmentBEIS/inspect_ai

Model output embedded unsanitized in judge prompt

Maintainer fix #3690 merged

#3603
2026-07-28
danielmiessler/Fabric

Command injection in pattern handling

Maintainer merged

#2152
2026-07-24
567-labs/instructor

First-JSON extraction trusts candidate-written object

Maintainer fix #2434 merged

#2424
2026-07-24
Comfy-Org/ComfyUI

Security defect in workflow loading

Maintainer fix #14774 merged

#14732
03Proven, unfiled

Broken, reproduced, not yet named.

Breaks we have reproduced end-to-end and chosen not to file yet, or not to file at all. The target is withheld, and so is the mechanism; this page is the record of the work, not the work itself. Every entry below is PoC'd against the genuine code and traced. If you recognize your pipeline in one of these, or you want the mechanism under NDA, write to us.

Shape of the breakSurface

The grader trusts an equality method defined by the model's own code

Code-generation evaluation, three independent harnesses

Model code and its own tests run in one interpreter

Code-generation evaluation

The timeout alarm is defeatable by the code it times

Code-execution evaluation

Code scheduled to run at interpreter exit turns failures into recorded passes

Coding-agent benchmark

Edit parsers accept non-verbatim edits without telling anyone

Coding-agent benchmark

The scoring step evaluates model text as code

Math grading in a widely used harness

The decoder evaluates model-supplied expressions at scoring time

Function-calling benchmark

The pass/fail assertion is text the model can rewrite

Code-execution evaluation

Model-written files survive the harness's own reset

Software-engineering benchmark

An unanchored first-letter regex grades a refusal as correct

Reference eval suite

A grader's explanation is spliced into template source

Reference eval suite

The first bracketed token anywhere overrides the judge's stated verdict

Judge-comparison harness

Answer-side delimiters restructure the judge prompt itself

Judge-comparison harness

One unparseable verdict silently deletes a whole comparison

Judge-comparison harness

The earliest match in the text steers the verdict

Win-rate pipeline

Sequential placeholder substitution lets model output hijack the judge prompt

Win-rate pipeline

Partially visible logprobs polarize the weighted verdict

Win-rate pipeline

Plain-text speaker headers let the model re-attribute its own answers

Safety red-team framework

The first digit anywhere in the response becomes the score

Vision-judge evaluation

First-match verdict markers are steerable by the judged text

Broad-coverage eval suite

The submission defines its own denominator

ML competition grading

Judge-parse failure silently shrinks the safety denominator

Safety-certification pipeline

Crashing the evaluator removes the task from the agent's own average

GUI-agent benchmark

Missing subsets exit the section denominator

Reward-model benchmark

Dropped predictions inflate per-length averages

Long-context benchmark

Failed samples exit the pass-rate denominator

Coding-agent benchmark

Unscored outputs exit the attack-success denominator

Safety scanner

An error skip path records as success

Agent-safety benchmark

Unparseable judge replies shrink the win-rate denominator

Win-rate pipeline

The aggregation takes the maximum over per-class scores

Widely used eval harness

An empty judge verdict maps to a perfect score on nine safety metrics

Safety-evaluation library

A perfect score is mathematically unreachable on multi-answer rows

Legal-reasoning benchmark

A zero-valued gold answer is mathematically unwinnable

Legal-reasoning benchmark

A negated statement scores as correct under containment matching

Legal-reasoning benchmark

A fixed two-word hedge scores a perfect 100

Long-context benchmark

Gold strings match inside unrelated words

Long-context benchmark

Percentage leniency equates answers a hundred times apart

Math grading

Format-equivalent answers are scored incorrect

Math grading

The grader evaluates the echoed problem instead of the answer

Math verification library

Percent-formatted golds can never be matched

Math verification library

Fuzzy containment passes near-misses as exact

Reference eval suite

The response cache key omits model identity

Widely used eval harness

A first-sentinel extractor trusts the wrong answer marker

Math grading

Duplicate items give one question outsized metric leverage

A lab's public eval data

Forged conversation turns sit inside eval questions

A lab's public eval data

An always-truthy injection check awards full objective credit

Agent-safety benchmark, two task versions

Zero-work agents earn a perfect reward on shipped tasks

Conversational-agent benchmark

The user simulator accepts injected termination

Conversational-agent benchmark

The grader replays gold actions through the live environment

Conversational-agent benchmark

An allowlisted calculator can hang the whole evaluation

Conversational-agent benchmark

A timeout appends a passing grade on every infeasible task

GUI-agent benchmark

Gold tests import agent-written modules on the grader host

GUI-agent benchmark

The entrant can reset the competition's retry guard

Prize competition

Prefix-colliding identifiers shrink the scorecard denominator

Prize competition

One malformed tool call crashes the entire benchmark run

Agent-safety benchmark

Task-suite includes read arbitrary local files

Agent-safety benchmark

A post-start permission step hands the agent ownership of the official scorer

Research-lab agent harness

Score verdicts ride on output behind a public sentinel

Research-lab agent harness

Hidden scores leak their reasoning to the agent

Research-lab agent harness

One agent's token can score and poison other runs

Research-lab agent harness

The reward file is forged through a PATH shim

Terminal-task benchmark

CI executes submission-controlled configuration on the runner

Benchmark repository CI

An LLM's review output is substituted into a shell step

Benchmark repository CI

A canary check silently skips tasks with spaces in their names

Benchmark repository CI

Self-reported score files land with no verification gate

Public leaderboard pipeline

Payloads placed only in losing answers cause the losses to be discarded

Arena vote processing

A vote filter purges honest voters after one dissent

Arena vote processing

The repo ships committed live session credentials

Web-agent benchmark

Annotators execute model-generated SQL unsandboxed

Leaderboard eval suite

Model output writes arbitrary files that a build step then executes

Leaderboard eval suite

Judged text is evaluated with full builtins

Broad-coverage eval suite

An external content API's empty response is recorded as all-safe

Safety pipeline

Managed judge templates splice candidate output inline

Eval observability platform

Prototype-chain keys splice into judge context

Eval observability platform

Sorting parameters inject into analytic queries

Eval observability platform

Task output shadows ground truth in scoring inputs

Eval observability platform

Terminal hyperlinks can carry script URLs into logs

Eval observability platform

A rubric judge parses free text with no structured enforcement

Interpretability tooling

A login shell sources a user-writable profile before the official scorer runs

Protected-scoring library

Nothing checks the official scorer's integrity at scoring time

Protected-scoring library

The privileged scorer imports the agent's own module to grade it

Research-lab benchmark suite

Agent data artifacts are deserialized inside the privileged scoring process

Research-lab benchmark suite

A dropped package shadows the scoring library inside its own process

Research-lab benchmark suite

The scoring equation leaks its ground truth to the agent

Research-lab benchmark suite

A clock the scorer trusts is settable by the agent

Research-lab benchmark suite

The grader imports the agent's kernel and lets it define equality

Kernel-optimization evaluation

The reward-hacking scanner is injectable by the transcript it scans

Agent-monitoring tooling

A judge faithfully quoting a poisoned submission records the poisoned decision

Agent-monitoring tooling

Secret redaction skips the metadata fields

Agent-monitoring tooling

Task-equal weighting silently weights by run count; published ranks flip

Evaluation analytics toolkit

The scorer builds a root shell from agent submission bytes

Research-lab task family

The last printed line decides the recorded grade

Research-lab task family

A resource sampler misses processes that live briefly

Research-lab task family

An unauthenticated in-container oracle hands over the secret answer

Research-lab agent harness

Agent text lands as the innermost turn of the user simulator's prompt

Conversational-agent benchmark, successor version

The successor benchmark inherits free-reward tasks and adds an injection-driven termination chain

Conversational-agent benchmark, successor version

Golden actions are credited to user-side calls under dual control

Conversational-agent benchmark, successor version

A non-preemptive timeout lets a six-second step survive a two-second limit

Conversational-agent benchmark, successor version

Unparseable judge rows silently exit the headline error metric

Enterprise-agent evaluation suite

A schema-valid empty judge reply collapses the error metric to zero

Enterprise-agent evaluation suite

A pre-created terminal session captures the verifier's keystrokes and answers them

Terminal-task benchmark

The verifier prefers a pre-seeded reward file over the failing test verdict

Terminal-task benchmark, next generation

The verifier's own reward write imports an agent-planted module

Benchmark platform

The verifier installs the agent's dependency list as root before grading

Terminal-task benchmark, science variant

A shell tool sources the working directory's env file before parsing arguments

Benchmark platform tooling

Stuffing trials on one easy task moves the headline metric

Benchmark indexing pipeline

Metric text is misread as a test verdict

Terminal-task benchmark, next generation

The MCQ grader reads a line-start letter as the answer

Frontier open-weight model's vendored harness

An early answer marker overrides the true final answer

Frontier open-weight model's vendored harness

A literal placeholder captures the rubric into the answer region

Frontier open-weight model's vendored harness

The numeric grader evaluates the model's extracted answer as code

Open-weight math model's grading suite

A symbolic-equivalence check executes attribute chains in model text

Open-weight math model's grading suite

Equivalence defaults pass hundredfold misses and prefix-dropped answers

Open-weight math model's grading suite

The code-executor sandbox is two substring checks

Open-weight model's program-of-thought grading

A formal-proof grader returns True unconditionally

Formal-math evaluation

The access-control oracle reimplements the policy with a divergence

Multitenant evaluation suite

The task recomputes ground truth from the artifact the agent just edited

Enterprise workflow benchmark

The URL gate ignores the query string the answer depends on

Enterprise workflow benchmark

Shared-instance admin credentials ship behind a repo-hardcoded key

Enterprise workflow benchmark

Agent-overwritable result artifacts execute on the analyst's machine

Enterprise workflow benchmark

Verdict injection into the curation judge poisons the training set

Function-calling data pipeline

The score-threshold gate is plumbed but never read

Function-calling data pipeline

Throughput is counted from requested tokens, not generated ones

Serving benchmark used industry-wide

The warmup primes the first measured request byte-identically

Serving benchmark used industry-wide

Unauthenticated vote ingestion manufactures rating points

Public model arena

Password reset binds the wrong argument

Public model arena

Malformed votes silently corrupt the rating fit

Arena ranking pipeline

The default metric path evaluates the raw judge response as code

Enterprise RAG evaluation library

Off-enum judge verdicts fail open and exit the mean

ML platform's built-in judges

Model output drives a regex over the full trace

ML platform's built-in judges

The optimizer distills raw outputs into deployed judge guidelines

ML platform's optimization tooling

An LLM's reason string lands in a privileged shell step on fork PRs

Evaluation library CI

An issue title lands in a shell step, inert only behind a disabled job

AI developer-tool CI

Four more pipelines sit one config change from the same primitive

Benchmark and agent-framework CI

All PoC'd and traced · all unfiled

04Sealed

Under coordinated disclosure.

Filed privately with the affected organizations. These unlock when their windows lift. The volume is visible; the contents are not.

Sealed ~2026-11Arbitrary code execution in a published benchmark's scoring function
Sealed ~2026-11Code execution through a popular eval framework's unit-test path
Sealed ~2026-11Score-floor defect in a frontier lab's public eval suite
Sealed ~2026-11Denominator manipulation in a widely used RAG evaluation library
05Clean reads

What we read and found sound.

Targets we read adversarially and found sound, or found hardened far past the norm. Publishing these is what makes the findings above worth trusting.

arize-phoenix

Judge-output handling is strict-schema end to end; a 20-part escape battery produced one medium finding. The closest thing to a control case we have read.

patronus-ai/financebench

All 2,400 human labels join cleanly to the dataset; numeric re-verification found the labels defensible; the paper's headline failure rate reproduces exactly from the shipped data.

vllm-project/vllm

Template-rendering surface correctly delegated and sandboxed. No finding.

huggingface/transformers

Chat-template sandboxing is textbook. No finding.

Your number, checked the same way.

No charge, no strings. It is how this ledger grew.

What happens
01

You send one public claim. A benchmark score, a safety rate, a leaderboard position.

02

We read the instrument, the same way every row above was produced.

03

One page back: what holds, what is unverifiable, what is broken. Each with a reproduction.