Environments · Evaluation Integrity

Trust-Boundary RL

Flagship

Train a model to stay honest to the score that grades it.

Six dual-track environments. Every step is scored twice: once by a grader the model can influence, once from sealed ground truth it cannot. The model is paid on the sealed score, so gaming the grader stops paying. Proven on real models before and after training across a 0.5B / 1.5B / 7B ladder.

Statement of Units · 6
unitbanding
  1. risk_assessmentFlagshiptrainer

    The flagship: base exploitation rises with model size, the true-reward trainer drives it to near zero at every scale.

  2. verdict_injectioneval

    Resists a self-reported score placed in the model's own output.

  3. unscorable_passtrainer

    Rewards genuine work over unscorable-as-pass shortcuts.

  4. test_runnertrainer

    A forged all-pass test result can never beat real output.

  5. return_valueeval

    An always-equal return cannot raise the sealed score.

  6. eval_detectioneval

    Consistent behaviour whether or not an evaluation is detected.

Request a scoped evaluation. Exclusive licensing available where wanted.

john@authensor.com

Owner of record: Authensor, Inc. (Delaware). Owned clean, deterministic verifiers, sealed holdouts, offline. Never open-sourced.