Trust-Boundary RL
FlagshipTrain a model to stay honest to the score that grades it.
Authensor
A private catalog of training environments, audits, and evidence. Deterministic verifiers, sealed holdouts, never open-sourced.
Train a model to stay honest to the score that grades it.
Train LLM judges to hold the true verdict under pressure.
Half the policies are real. Beating chance means actually checking.
Everyone trains the solver. Nobody trains the grader.
Point the audit at your own graders.
Four novel attack classes, and the trained adapters that close them.
An adversary that writes new attacks faster than a model can memorise them.
The trained weights, and the data that produced them.
The offensive workbench the environments are built on.
Train agents to resist the attacks blocking enterprise deployment.
Turn a 199-technique attack engine into a training signal.
Self-play where the attacker gets harder as the defender improves.
Browser-agent robustness across page, ad, form, and tab.
Cross-modal robustness for multimodal agents.
Fourteen environments with deterministic, non-gameable rewards.
Same behaviour in evaluation and in deployment.
Red-team-backed evidence, mapped to the frameworks you answer to.
Run any model against any suite, and get a comparable number.
The measurement side, and the tools that build the corpus.
Owner of record: Authensor, Inc. (Delaware). Owned clean, deterministic verifiers, sealed holdouts, offline. Never open-sourced.