Why Are We Building the Car Before the Train?

America is racing toward general-purpose AI agents with open-ended agency. China is putting its industrial weight behind AI inside defined systems and workflows. Two centuries of transportation engineering suggest we are handing over the steering wheel before we have laid the track.

Subscribe
AgentsAIInfrastructurePolicy
Listen to this post
AI Summary American AI development is organized around general-purpose agents that receive objectives, inspect environments, build plans, choose tools, and act with less human oversight, while China's industrial policy directs AI into specific manufacturing processes with defined workflows, datasets, and application scenarios. …
  • American AI development is organized around general-purpose agents that receive objectives, inspect environments, build plans, choose tools, and act with less human oversight, while China's industrial policy directs AI into specific manufacturing processes with defined workflows, datasets, and application scenarios.
  • Using transportation data from 2000 through 2014, cars and light trucks had 6.53 passenger fatalities per billion passenger-miles compared to roughly 0.36 for commuter and intercity rail—a difference of roughly eighteenfold—because rail safety is a property of the system around the train, not the operator's intelligence.
  • Risk in AI systems tracks the product of capability, permissions, duration, and reachable state space, and general-purpose agents run an open loop where they write the workflow as they go rather than executing a fixed workflow, making errors stateful as wrong decisions at step 27 shape the environment encountered at step 28.
  • Mature engineering in aviation, nuclear plants, payment networks, and operating systems achieves safety by restricting systems until their behavior is understood and then widening the envelope, but AI is tempted to run this backwards by releasing machines that navigate arbitrary digital environments before anyone has characterized what happens when they operate autonomously.
  • The better architecture separates cognitive freedom from operational freedom by letting the model reason with wide freedom while execution passes through a separate layer that validates whether proposed actions fall inside acceptable bounds—a railway that can lay its own track but cannot move until the track is inspected.

A divergence is emerging in how the United States and China approach artificial intelligence. American companies build plenty of tightly constrained systems, and Chinese labs ship increasingly general models and agents, so neither country has settled on a single architecture. The centers of gravity still sit in different places, and the difference is large enough to matter.

American frontier development is organized around the agent: a system that receives an objective, inspects its environment, builds a plan, chooses tools, acts, observes the results and revises as circumstances change.

  • OpenAI describes agents as systems that complete multi-step work with tools while carrying context across the steps, and its computer-using agent operates graphical interfaces built for humans rather than waiting for every action to be exposed through an API.12
  • Anthropic describes the same transition, from question-and-answer toward agents that execute code, manage files and work across applications, and names the cost plainly: agents act with less human oversight, which leaves more room to misread intent and take actions with unintended consequences.3

China's industrial policy puts its emphasis elsewhere.

The AI + Manufacturing implementation opinions, issued in January 2026 by the Ministry of Industry and Information Technology with seven other agencies, direct AI into specific manufacturing processes, industrial systems and operating scenarios.

The 2027 targets are concrete:

  • three to five general-purpose models deployed deeply in manufacturing
  • 1,000 high-level industrial agents
  • 100 high-quality industrial datasets
  • 500 representative application scenarios.4

An August 2026 MIIT program follows up by cultivating a layer of AI application service providers — firms that understand an industry's problems, integrate systems, run delivery and handle safety and compliance.5

Important

Both paths lead toward more capable machine intelligence, but they rest on different philosophies. One invests in making the driver more capable. The other spends far more of its effort building the tracks. Transportation has a long record of what each choice produces.

Cars vs Trains

Cars Have Agency. Trains Have Constraints.

A driver chooses where to go. They can change lanes, turn onto another road, accelerate, brake, take a shortcut, miss an exit, pull onto private property or leave the road entirely. That flexibility is enormously valuable, and it produces an enormous state space.

A train operator has far less freedom. The train cannot turn left. It cannot decide the neighboring field looks like a faster route. Its movement is bounded by rails, switches, signals, schedules, dispatchers, operating rules and a growing layer of automated protection. Operating a train can still demand real skill; what has been deliberately removed is agency over most of the environment.

The safety gap is stark.

Using Bureau of Transportation Statistics data for 2000 through 2014, the American Public Transportation Association calculated 6.53 passenger fatalities per billion passenger-miles for cars and light trucks, against roughly 0.36 for commuter and intercity rail — a difference of roughly eighteenfold.67 Tracks alone do not account for all of it. Cars and trains run in different environments, with different infrastructure, loads, procedures and exposure. That is exactly what makes the comparison instructive: rail safety is not a property of the train. It is a property of the system around the train.

A train is not safer because its operator is smarter. It is safer because the system has removed most of the ways the operator could be wrong.

Anyone who has built reliable software recognizes the principle. Type systems, permissions, sandboxes, transaction boundaries, rate limits, schemas, memory protection, network segmentation, circuit breakers and rollback all do the same job. Reliability comes from narrowing what each component can touch, not from granting every component access to everything and trusting it to reason well.10

General-purpose agents are drifting toward the second model.

Task-oriented workflows vs General purpose AI agents

Intelligence and Agency Are Not the Same Thing

Somewhere in the development of modern AI, intelligence and agency began to merge. As models reason better, they get more tools. As they use tools better, they get more permissions. As they plan better, they run longer sequences without intervention.

The implied progression is

flowchart LR
    A[More Intelligence] --> B[More Tools]
    B --> C[More Permissions]
    C --> D[More Autonomy]

and no law of computer science requires it. A system can reason brilliantly while holding almost no authority over the world, and a mediocre system becomes dangerous once it holds enough privilege and enough chances to act.

Risk tracks the product of capability × permissions × duration × reachable state space, and the last term gets the least attention.

  • A task-oriented system runs a fixed shape:
flowchart LR
    objective --> approved_tools[approved tools]
    approved_tools --> defined_workflow[defined workflow]
    defined_workflow --> validation
    validation --> execution
    execution --> output
  • A general-purpose agent runs an open loop:
flowchart LR
    A[Objective] --> B[Inspect Environment]
    B --> C[Construct Plan]
    C --> D[Select Tools]
    D --> E[Act]
    E --> F[Observe Consequences]
    F --> G[Revise Plan]
    G --> H[Continue]
    H --> B

The second is far more powerful and far harder to characterize, because the agent is no longer executing a workflow. It is writing the workflow as it goes. That is the AI equivalent of leaving the rails.

The hard problem may not be making AI intelligent enough to drive. It may be explaining why intelligence should automatically earn the right to steer.

The Reachable State Space Explodes

Conventional software reliability leans heavily on knowing which states a system can occupy. A database transaction has defined states, a finite-state machine has enumerated transitions, an API exposes specific operations, and a permission system spells out what each actor may do. Engineering gets much harder once the set of reachable states becomes effectively unbounded.

A single model response can be evaluated. A thousand-step interaction with a live environment is a different object, because the output of step 27 becomes part of the world encountered at step 28. A misreading shapes a tool call, the tool call changes the environment, and the changed environment becomes evidence for the next decision. Errors become stateful.

The system does not just produce a wrong answer; it produces a wrong world and keeps reasoning inside it.

The frontier labs see this clearly.

  • Anthropic's own framework flags reduced oversight, unintended actions and prompt injection as risks that intensify as agents take on more consequential work.3
  • OpenAI's guidance for builders stresses guardrails, scoped tool access, approvals and human intervention even as it pushes toward more independent multi-step execution.8

Recognition of the problem is not in doubt. The open question is whether the product trajectory is moving toward agency faster than the industry's understanding of how agency fails.

Mature Engineering Starts With Constraints

The pattern is consistent: restrict the system until its behavior is understood, then widen the envelope.

  • Aviation did not become extraordinarily safe by giving pilots more latitude. It got there through controlled airspace, standard procedures, checklists, redundant systems, flight envelopes, air-traffic control, certification, instrumentation and relentless accident investigation.
  • Nuclear plants wrap powerful processes in containment, interlocks and independent layers of protection.
  • Payment networks impose transaction limits, settlement rules, identity requirements and fraud controls.
  • Operating systems isolate programs in separate processes instead of letting any program write anywhere in memory.

AI is tempted to run it backwards — to release machines that navigate arbitrary digital environments before anyone has characterized what happens when they operate autonomously inside those environments for long stretches.

Warning

We are building the car while inventing the driver, the roads, the traffic laws, the insurance market and the crash standards in parallel, and removing the person from the passenger seat as we go.

China Is Asking a Different Question

China is pursuing frontier models, robotics and autonomous systems, and nothing in its policy rejects general intelligence.12 Its published industrial strategy, though, starts from scenarios: manufacturing lines, supply chains, agriculture, industrial software, equipment maintenance, production optimization, telecommunications. The sequence runs from a defined operating environment, to data assembled around it, to specialized agents introduced into it, to intelligence integrated into a system that already exists.459

That sequence places much of the intelligence in the environment around the model. The agent does not have to discover the whole operating system for itself, because the workflow, the permissions, the available tools, the schema and the process already carry knowledge. The tracks embody engineering decisions made before the AI arrives.

The American frontier-agent program asks how intelligent the driver can become. The scenario-first program asks how intelligent the transportation system can become. Those questions produce very different architectures.

The Answer Is Neither the Car Nor the Train

Tracks have an obvious weakness: they reach the destinations their designers anticipated and nothing else. Destinations that do not exist yet are out of reach, and that is precisely why general-purpose agents are so compelling. The ability to walk into an unfamiliar environment and work out what to do is worth a great deal, so the goal cannot be to lock AI into fixed workflows forever.

The better architecture is a railway that can lay its own track but cannot move until the track is inspected.

The model reasons with wide freedom — exploring options, constructing plans, finding tools, simulating consequences — while execution passes through a separate layer that decides whether the proposed actions fall inside acceptable bounds:11

flowchart LR
    A[Reason freely] --> B[Propose route]
    B --> C[Validate route]
    C --> D[Authorize actions]
    D --> E[Execute within bounds]
    E --> F[Monitor]
    F --> G{Deviation?}
    G -->|Yes| H[Stop or roll back]
    G -->|No| F

This separates cognitive freedom from operational freedom, and the separation carries most of the weight. An AI does not need authority to wire $10 million because it understands why wiring $10 million might help. It does not need an unrestricted shell because it writes excellent code. It does not need access to every corporate system because it grasps the company's goals. Intelligence can stay expansive while authority stays deliberately narrow.

The future may belong not to the AI with the most agency, but to the system that extracts the most intelligence while granting the least unnecessary agency.

None of this rules out autonomous AI. It calls for progressive autonomy: build the reasoning, then the instrumentation, the evaluations and the containment; learn the failure modes; widen the envelope; widen it again. Some applications will eventually want something close to the autonomous car.

Most need the railway first.

Railways learned this lesson in the nineteenth century, and the device they built for it was the interlocking — a frame of levers mechanically linked so that a signalman physically could not clear a signal onto a route whose switches were set against it. The signalman stayed fully in charge of the plan. The machine simply refused to let two trains be sent onto the same stretch of track.

That is the missing layer in agentic AI: not a smarter driver, but an interlocking between what the model decides and what the world lets it do.

Footnotes

  1. Agents — OpenAI developer documentation — OpenAI's own definition of the architecture this piece is describing: systems that plan and complete multi-step tasks using tools while maintaining context across the steps. The open loop is not a critic's characterization — it is how the capability is documented by the people shipping it. https://platform.openai.com/docs/guides/agents ↩
  2. Computer-Using Agent — OpenAI, January 2025 — The specific escalation that matters to the argument: an agent that operates graphical interfaces built for humans and adapts as the task unfolds, rather than waiting for every action to be exposed through an API. An API is a declared surface with a known set of operations. A screen is not, which is what makes the reachable state space hard to bound. https://openai.com/index/computer-using-agent/ ↩
  3. Trustworthy agents in practice — Anthropic, April 9, 2026 — The cost named by the lab shipping the capability, in its own words: models that "can write and execute code, manage files, and complete tasks that span multiple applications," and agents that "act with less human oversight, so there is more room for them to misread users' intent" and take actions with unintended consequences. The same page treats prompt injection as attacks that try to trick models into taking costly actions they otherwise would not. Recognition of the failure mode is not the disputed part of this argument. https://www.anthropic.com/research/trustworthy-agents ↩
  4. Implementation Opinions on the "AI + Manufacturing" Special Initiative — MIIT and seven other agencies, January 7, 2026 (CSET translation) — The scenario-first strategy in its own targets, all for 2027: three to five general-purpose models deployed deeply in manufacturing, 1,000 high-level industrial agents, 100 high-quality industrial datasets and 500 typical application scenarios. Issued jointly by MIIT with the Office of the Central Cyberspace Affairs Commission, the NDRC, the Ministry of Education, the Ministry of Commerce, SASAC, the State Administration for Market Regulation and the National Data Administration — eight bodies, which is itself a statement about where the effort sits. https://cset.georgetown.edu/publication/china-ai-plus-manufacturing-initiative-opinions/ ↩
  5. China Eyes 3,000 AI Service Providers by 2027 — Science and Technology Daily, September 15, 2026 — The service-provider layer this sentence refers to: a national resource pool of AI application service providers expected to exceed 2,000 by the end of 2026 and reach at least 3,000 by the end of 2027, built to give industries consulting, deployment, operations and safety governance. This is the part of the strategy that has no clean American analogue — the state is funding the integrators rather than the models. https://www.stdaily.com/web/English/2026-09/15/content_580446.html ↩
  6. The Hidden Traffic Safety Solution: Public Transportation — American Public Transportation Association, September 2016 — Table ES-2, on Bureau of Transportation Statistics data for 2000 through 2014: 6.53 passenger fatalities per billion passenger-miles for car or light-truck occupants. https://www.apta.com/wp-content/uploads/Resources/resources/reportsandpublications/Documents/APTA-Hidden-Traffic-Safety-Solution-Public-Transportation.pdf ↩
  7. Testimony of Paul P. Skoutelas, President and CEO, APTA — House Committee on Transportation and Infrastructure, February 15, 2018 — The rail figure set directly against the automobile one: 0.36 passenger fatalities per billion passenger-miles for commuter and intercity rail against 6.53 for automobiles and light trucks. APTA states the ratio as rail being eighteen times safer for passengers, which is the multiple used here. https://transportation.house.gov/uploadedfiles/2018-02-15_-_skoutelas_testimony.pdf ↩
  8. A Practical Guide to Building Agents — OpenAI, 2025 — The guidance that sits alongside the capability: agents that independently execute workflows, with guardrails, tool safeguards, scoped access, approvals and human intervention presented as the conditions for predictable operation. The tension this piece is pointing at lives inside a single document — the containment advice and the autonomy push are published together. https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf ↩
  9. Opinions of the State Council on Deepening the Implementation of the "Artificial Intelligence+" Initiative — August 26, 2025 (CSET translation) — The parent document for the scenario-first sequence, directing AI across industrial design, production, services, operations, supply chains, agriculture and other defined sectors. Its phased targets: penetration of next-generation intelligent terminals and agents above 70% in key domains by 2027, above 90% by 2030. Official English release at https://english.www.gov.cn/policies/latestreleases/202508/27/content_WS68ae7976c6d0868f4e8f51a0.html https://cset.georgetown.edu/publication/china-ai-plus-opinions-2025/ ↩
  10. The Honor System Does Not Scale — The same claim made about coordination rather than containment: agent trust is not a problem better prompts can solve, because it is a property of the infrastructure rather than of the participants. Narrowing what a component can touch and establishing what a counterparty can be held to are two faces of one argument — reliability is built outside the model or it is not built. ↩
  11. The Room Scales — What a validation layer has to survive. That piece argues for the one auditing test that does not degrade as the systems being audited get more capable, which is the requirement any route-inspection layer inherits — a check that a smarter agent can talk its way past is not a check, it is a formality. ↩
  12. What Export Controls Actually Accelerated — Why the scenario-first emphasis should not be read as a capability ceiling. That piece argues export controls pushed Chinese development toward the kind of efficiency that matters most, which is the same observation from the other direction: a constraint on one axis redirects effort rather than stopping it. ↩
Back to the Journal