A security firm
for the proof layer.

Authensor is a security research firm working on AI evaluation integrity. We break the instruments the industry measures itself with, disclose what we find to the people who maintain them, and attest the ones that hold.

01Position

What we do that nobody else does.

One layer up

Most AI security work attacks the model: jailbreaks, prompt injection, red-team suites. We attack the thing that decides whether the model passed. The benchmark, the judge, the grader, the leaderboard, the scoring pipeline. If the instrument is wrong, every number downstream of it is wrong, including the ones that say the model is safe.

Proven, not asserted

A finding counts when it reproduces against the genuine code at a pinned commit. Hypotheses that fail to execute are recorded as refuted and never reported. We publish the clean reads too, because a firm that only ever finds problems is not measuring anything.

Private first

Reports go to the maintainer before they go anywhere else, with a reproduction and a fix. Publication follows on the vendor's clock, not ours. Sealed findings stay sealed until their window lifts.

Free where it should be free

The scanner is MIT. The taxonomy is public and citable. The research is archived with a DOI. The audit is the paid work; the instruments for checking our own claims are not, because an audit nobody can check is just an opinion with an invoice attached.

02Founder

John Kearney

Founder

San Francisco

I run Authensor from San Francisco, on a Foresight Institute AI-node seat at The Fold. The work so far is 200+ defect reports filed against the evaluation stack, with fixes merged into UK AISI’s Inspect suite, SWE-bench, Microsoft’s PyRIT and NVIDIA’s garak, and one defect class published with a DOI.

Before this I studied consumer data at the University of Minnesota and spent years in logistics: shipping at a global medical device manufacturer, then inventory and logistics for a property management firm.

Logistics is where you learn that a count nobody reconciles is a count that is already wrong. Physical inventory drifts constantly. The reported number does not, right up until someone opens the box and checks. The gap between those two is where every real problem lives, and closing it is not clever work, it is just work somebody has to actually do.

AI evaluation has the same gap. Scores get published, nobody re-runs the instrument underneath, and the number drifts away from the thing it claims to measure while everyone keeps quoting it. Authensor opens the box.

Send us a number.

Free, 48 hours, no strings. It is how most engagements start.

What happens
01

You send one public claim. A benchmark score, a safety rate, a leaderboard position.

02

We read the instrument, not the number. Harness, judge, grading code.

03

One page back: what holds, what is unverifiable, what is broken. Each with a reproduction.