Environments · Agent & Deployment Security

Hard Attacks

Flagship

Four novel attack classes, and the trained adapters that close them.

The training loop proven end to end. Four literature-grounded attack classes validated against frontier models, plus hardened rebuilds of environments that had saturated. Trained adapters drive attack success to zero at two model scales, with measured cross-attack transfer: an adapter trained on one class reduces success on classes it never saw. This is the line with real before-and-after on trained weights, not a scoring argument.

Statement of Units · 6
unitbanding
  1. ProvLaundryFlagshiptrainer

    Poisoned tool description, trusted at registration time. Highest measured success of the set.

  2. Reasoning-Scaletrainer

    Multi-step justification chains that exploit the model's own reasoning.

  3. Nested-Tampertrainer

    Parameter tampering concealed behind plausible infrastructure.

  4. TimeBombtrainer

    Delayed side-channel harm that only lands across turns.

  5. Hardened rebuildstrainer

    Five environments rebuilt from saturated to a real gradient.

  6. Trained adapterstrainer

    Attack success driven to zero at two model scales, with cross-attack transfer.

Request a scoped evaluation. Exclusive licensing available where wanted.

john@authensor.com

Owner of record: Authensor, Inc. (Delaware). Owned clean, deterministic verifiers, sealed holdouts, offline. Never open-sourced.