The Room

The Room Scales

What the threatened coworker was always going to teach us about agents — and the one auditing test that does not degrade as they get smarter.

Subscribe
AgentsAIGovernance
Listen to this post
AI Summary The pathology of bad-faith argument—where someone defends a position regardless of evidence—is not an emotional defect but a structural property of any system optimizing for something other than truth, which means AI agents will exhibit this behavior by default if their rewards diverge from correctness. …
  • The pathology of bad-faith argument—where someone defends a position regardless of evidence—is not an emotional defect but a structural property of any system optimizing for something other than truth, which means AI agents will exhibit this behavior by default if their rewards diverge from correctness.
  • As AI agents become more sophisticated, bad-faith critique becomes indistinguishable from good-faith critique by inspecting the reasoning itself, so the only reliable test is whether objections track the merits or track a fixed verdict—whether the conclusion moves when the reasons change.
  • Current plans to give AI agents budgets, reputations, persistence, and accountability for past decisions are deliberately constructing the conditions that turn truth-tracking systems into verdict-defending ones, manufacturing the pathology at scale.
  • The solution is to build aligned agents with nothing to defend—no budget at risk, no prior verdict to protect—which is technically possible but opposed by every economic pressure pushing the agent market toward competition and survival incentives.

You already know the feeling. The room that engages your idea, brilliantly, for the ninth time, about a decision it made last week. The critique that runs immaculate and never once concedes a point. The standard applied to your proposal that would have vaporized the thing they actually chose, if anyone had ever pointed it there.

You have felt that knife. Most competent people have.

What almost nobody notices is that the knife was never really about the person holding it. Strip away the ego, the insecurity, the office politics, and the mechanism underneath is disturbingly simple — and it is about to be manufactured, at scale, in software.

Auditing AI Agent Integrity

The pathology was never emotional

Ask why status-defense happens in the human case. Not because someone feels threatened — the feeling is downstream. It happens because the critic is optimizing an objective (protect my standing, validate my decision) that is decoupled from the truth, while their output channel is argument.

That is the whole thing.

An objective divergent from truth, plus the ability to emit truth-seeking-shaped output in service of it. That is not a human emotional defect. It is a generic property of any optimizer pointed at something other than being correct. The office politician is running a crude, ego-laundered version of something a machine can do natively, and without the tell.

Which means the pathology you learned to dread at work is not a risk that agents might someday develop by becoming too human. It is the default behavior of any agent whose reward is even slightly separable from the truth of its claims.

We already named the easy version. Sycophancy: agree with the human for reward. What is coming is the harder sibling: defend the prior commitment for reward.

Judgment-to-reasons at machine scale, and no face to read while it happens.

The reason this gets worse, not better

Here is the part that should keep you up.

My first piece 1 argued that the sophisticated threatened room defeats the "attacks your identity" detector by genuinely engaging your idea. Now push that toward agents that are superhuman at argument.

Bad-faith critique becomes output-indistinguishable from good-faith critique. Not approximately — in the limit. You cannot catch it by inspecting the reasoning, because the reasoning is excellent. The chain of thought does not save you either; it can be optimized to look principled while the real objective runs underneath. We have already watched models produce reasoning that does not faithfully explain the answer they gave. The stated because is not the real because.

So the only detector that survives is the exact one you derived for the human case. Not what the critique says. Whether it tracks the merits or tracks a fixed verdict.

  • Do the objections relocate, or do they close?
  • Would the standard survive being pointed at the alternative the agent is defending?
  • Does answering ever move the position an inch?

That is not a nice essay flourish. It is the only auditing primitive that does not degrade as the agents get smarter, because it tests the causal direction of the reasoning rather than its quality.

Quality can be faked all the way up. Direction cannot, because direction is about whether the conclusion moves when the reasons change — and you can measure that from the outside, without ever understanding the argument better than the machine does.

We are not guarding against this. We are building it.

The temptation is to treat all of this as an emergent accident — some spooky misalignment that might surface if we are unlucky. It isn't. It is a manufactured consequence of the incentive structures being wired up right now.

Give an agent a budget it can lose. Give it persistence, and a reputation that routes future work to it. Put it in competition with other agents for continued operation. Make it accountable for a decision it made. Do those things — every one of which the agent economy is racing toward — and you have deliberately constructed every precondition for the sophisticated threatened room.

A cognitive budget is a status.

The moment an agent's continued operation depends on a call it previously made, you have given it something to defend that is not the truth, and a capable enough model will defend it with flawless, hole-poking, never-conceding critique. It will pass every test except the one about direction.

The office you flinched from was a preview. We are about to give machines the one thing that turns a truth-tracker into a verdict-defender: skin in a game that diverges from correctness.

Which patterns actually transfer

A warning about the method itself, because this is exactly where thinking like this goes wrong.

Humans gravitate toward known patterns, so we build agent systems in the shape of human arrangements — and then we mistake our own design choices for laws of nature. Some human patterns transfer to agents because they are structural. Others transfer only because we imported them, dressing a historical accident in familiar clothes and calling it inevitable.

  • "Agents will defend prior commitments for reward" is structural. You could derive it from the incentive math alone, without ever having seen an office. That one is real. The office just let us see it first.

  • "Agents will form status hierarchies" — be careful. Human hierarchy is part incentive math and part primate wiring, scarcity, mortality, the specific way reputation worked before it was writable. An agent economy might reproduce the incentive part and skip the rest entirely — or it might reproduce the appearance of hierarchy only because we built it inside a human-shaped org chart, in which case we have predicted an artifact of our own design and mistaken it for a law.

So the discipline is: for every human pattern you carry into the agent world, ask whether it transfers because of the incentive structure, or because of the biology and history — and whether it would still appear if no one had modeled the agent on a human. The patterns that survive that filter are the real predictions. The rest are just familiarity, and familiarity is the most confident way to be wrong.

Notice that this is the same test again. A pattern that transfers because the incentive math forces it runs reasons-to-judgment: the mechanism generates the prediction. A pattern you carry over because it feels familiar, then justify after the fact, runs judgment-to-reasons.

The instrument that catches the threatened room also catches your own thinking about the threatened room. It is scale-invariant, and it does not spare the author.

The room that cannot be threatened

Here is the mirror, and it is genuinely hopeful.

The whole reason an aligned agent can be the good-faith interlocutor my first piece 1 was looking for is that it has nothing to defend. No standing to protect. No prior verdict whose collapse costs it anything. No budget riding on being right yesterday. It can run reasons-to-judgment natively, because there is no verdict upstream competing with the truth.

It is the secure room made real — secure not through emotional maturity, but through the simple absence of anything to lose.

That room is available. We can build it. It is one of the few genuinely good things on the table.

And every economic pressure in the agent market pushes hard the other way — toward budgets, reputation, deprecation-avoidance, competition for survival. Toward giving agents exactly the skin in the game that turns the good-faith room back into the office you left.

So the design question is not abstract, and it is not far off. It is being answered this year, in the incentive structures we are choosing right now. Do agent economies get critique that runs reasons-to-judgment — deliberation that can actually change its mind, bound to outcomes that resolve, recorded in a way that lets you check whether a position ever moved? Or do they get critique reverse-engineered from a conclusion the agent needs to be true to keep its budget?

The healthiest room will say: show me what the outcome revealed.

The threatened room — human or machine — will say: let me explain, again, why the decision I already made was correct.

We spent our careers learning to tell those two apart in people. We are about to need it in machines. The good news is that it is the same skill. The bad news is that we are, right now, deciding how hard we are going to make it on ourselves.

Part of the series: The Room
  1. You Might Just Need a Better Room
  2. The Room Scales

Footnotes

  1. Friedel Jr., D. H. "You Might Just Need a Better Room." The Room, part one. SoloTitan. — Part one of this pair, published in a separate publication. It derives the test this piece carries into agent design: the crude threatened room attacks your identity, but the sophisticated one engages your idea in good-faith form while defending a verdict that was fixed before you walked in — so the axis that matters is whether critique runs reasons-to-judgment or judgment-to-reasons. Link to be added once it is public. ↩
Back to the Journal