The Interlock

How small per-step error rates become large harms over long agent runs, and how much a validation layer between decision and action changes the outcome. Illustrative model, not a forecast: the inputs are levers, not measured rates.

The agent

Chance any single action is wrong. Rises with reachable state space.
Actions taken without a human resetting the run.
Blast radius of one uncaught action. Each level is ×10.
After an uncaught error the agent reasons inside a wrong world: each one raises the error rate by this much.

The system around it

Share of wrong actions blocked before execution. Blocked errors never touch the world.
Per-step chance a deviation, once it exists, is noticed and the run stopped.
Share of harm recoverable by rollback.
P(harmful action in run)
Expected harm per run
Worst 1% of runs (P99)
Uncaught errors per run

Cumulative chance of a harmful action

Share of simulated runs that have executed at least one wrong action by step t
This configurationSame agent, no interlock or monitor

Distribution of harm across runs

Runs with any harm, binned by total harm (log scale, harm units). Harm-free runs are excluded.
This configurationSame agent, no interlock or monitor

Model: each step, the agent proposes a wrong action with probability p. The interlock blocks it with probability c. An unblocked error executes, causes harm of 10^scope × lognormal(σ=1) × (1 − reversibility), and multiplies p by (1 + compounding), capped at 50%. Once any error has executed, the monitor halts the run with probability m per step. Each view is 2,000 Monte Carlo runs on a fixed seed, so moving one lever changes only what that lever controls. Values are in arbitrary harm units; the shape of the curves is the finding, not the numbers.