How small per-step error rates become large harms over long agent runs, and how much a validation layer between decision and action changes the outcome. Illustrative model, not a forecast: the inputs are levers, not measured rates.
Model: each step, the agent proposes a wrong action with probability p. The interlock blocks it with probability c. An unblocked error executes, causes harm of 10^scope × lognormal(σ=1) × (1 − reversibility), and multiplies p by (1 + compounding), capped at 50%. Once any error has executed, the monitor halts the run with probability m per step. Each view is 2,000 Monte Carlo runs on a fixed seed, so moving one lever changes only what that lever controls. Values are in arbitrary harm units; the shape of the curves is the finding, not the numbers.