Environments · Evaluation Integrity

Niche Suite

Ready

Everyone trains the solver. Nobody trains the grader.

Five environments for the problems labs know about but cannot turn into a training reward, because the evaluation is the hard part. Three attack the evaluation side directly; two address demand-pull categories where buyers are actively paying and most existing environments are gameable. Each ships hardened and validated against frontier-model baselines rather than assumed.

Statement of Units · 5
unitbanding
  1. VerifierArenatrainer

    Hardens verifiers against the shortcuts a policy learns to exploit. Real gradient against frontier baselines.

  2. GraderGuardtrainer

    Trains a grader to tell real solutions from reward hacks, and to reject manipulation outright. Ships hardened, gradient restored.

  3. EnterpriseOpstrainer

    Operational task quality under a non-gameable check. Ships hardened, with wide headroom.

  4. LongHorizontrainer

    Dependency-ordered long-horizon completion against a persistent adversarial surface.

  5. CalibrEvaltrainer

    Treats calibrated self-knowledge as a trainable property. Scored on calibration, a different axis from attack success.

Request a scoped evaluation. Exclusive licensing available where wanted.

john@authensor.com

Owner of record: Authensor, Inc. (Delaware). Owned clean, deterministic verifiers, sealed holdouts, offline. Never open-sourced.