Eminence · Article 06 · September 30, 2026
The factory is the system
Design the handoffs, checks and learning around the model.
When people talk about building an AI product, they usually point to the model. That is a little like pointing to a machine on a factory floor and calling it the factory.
The machine matters. But the customer experiences the whole line: what comes in, what gets transformed, where work waits, how errors are found, who can stop production, and what happens when demand changes. A brilliant machine in a badly designed line can make defects faster.
Manufacturing has spent more than a century learning this the hard way. Its history offers a useful way to design AI systems, provided we transfer the operating principles rather than pretend that answers are identical to car parts.
First, make the parts fit
Before the moving assembly line, manufacturers were already learning to make parts interchangeable. Ford's contribution in 1913 was to combine standardized parts, divided work and the movement of work between stations. The gain came from designing the relationships between steps, not merely making each worker faster. The Henry Ford museum records the progression from magneto experiments to chassis assembly and then continuous vehicle lines. Source: The Henry Ford[1].
For an AI system, the equivalent of an interchangeable part is a clear handoff. A request should arrive with a defined purpose and allowed data. Retrieval should return material with source and date. A model should produce an answer in a form that the next step can check. A tool should return either an observable result or a specific failure. When these interfaces are vague, every change becomes a custom repair.
This is the first tenet of system design: design the connections, not only the components. A better model can improve a station. A better handoff can improve every model you might use later.
Then, measure the process rather than admiring the output
In the 1920s, Walter Shewhart's work on control charts gave manufacturers a way to distinguish ordinary process variation from signals that something had changed. The point was to learn about the process while it was running, not discover its failures only in a pile of finished goods. Source: American Society for Quality[2].
AI systems have variation too. The same question can produce different answers. The inputs change. The documents change. Users ask for things the designer never expected. One impressive demo tells us almost nothing about the range of outcomes.
So measure the work by type. How often is a cited claim actually supported? What happens when the source is missing or contradictory? How often does a tool fail, and does the system retry safely? Which errors reach the user? Google researchers have long argued that production machine learning depends on tests and monitoring around the model, not only model quality in isolation. Source: Google Research[3].
The second tenet: treat quality as a property of a repeatable process. An average score can hide a dangerous minority of cases. Look for patterns, especially at the edges.
Build the stop into the line
Toyota added another crucial idea. Its production system pairs just-in-time flow with jidoka: detect an abnormality and stop the process before defects travel downstream. That idea began in Sakichi Toyoda's looms, which could stop when thread broke. The machine did not need a person staring at it every second; it did need a way to make a problem visible. Sources: Toyota's loom history[4] and Toyota Production System history[5].
An AI workflow needs its own stop points. A research assistant should say when the evidence is absent. An agent should stop before an irreversible action whose authority it cannot verify. A document generator should flag a broken citation before polishing the final PDF. Those are design choices, not apologies for imperfect technology.
The third tenet: catch defects where they appear. If the only quality check is a human cleaning up the final answer, the system has already carried the defect through every later step.
Find the constraint, then decide what to automate
A factory can have ten busy stations and still be limited by the one that cannot keep up. AI teams make the same mistake when they optimize the model call while a human approval queue, poor source data or a fragile integration controls the actual pace of work. Map the flow from the user's request to the delivered outcome. Count waiting, rework and rejected outputs as part of the cycle. Then improve the constraint that governs the result.
This is where Elon Musk belongs in the manufacturing timeline. In Tesla's 2016 master plan, he argued that the factory itself should be designed as a product: the “machine that makes the machine.” That is a valuable systems insight. Tesla's subsequent Model 3 ramp also shows its limit. In a 2018 filing, the company said it had temporarily reduced automation in some processes and added semi-automated or manual work while addressing bottlenecks. Sources: Tesla's 2016 plan[6] and 2018 Form 10-Q[7].
For AI, “automate everything” is no more of a strategy than “put a robot at every station.” Automate the stable, well-understood step when it improves the whole flow. Keep people where judgment, exception handling or authority are the real work. Revisit the boundary when the process changes.
The fourth tenet: optimize the system's outcome, not its automation percentage.
Give the system a way to learn
The deepest shift in systems thinking is to stop seeing failures as isolated blemishes. A failed output can reveal a bad source, an incentive to rush, a missing permission, an unclear goal or a feedback loop that rewards the wrong thing. Donella Meadows argued that information flows, rules and goals can matter more than tuning a parameter. Deming similarly treated variation, knowledge, people and the whole system as connected parts of management. Sources: Donella Meadows[8] and The Deming Institute[9].
Start by collecting the right feedback. Keep the request, the source version, the system's response, tool failures, human corrections and, when it can be observed, the outcome the user actually got. Collect only what the work needs, with appropriate permission and retention. Otherwise, the team is left with a thumbs-up, a complaint or an impressive dashboard, none of which explains what to change.
Turn those observations into AI evaluations: repeatable cases that test whether the system completes the task, supports its claims, handles missing evidence and stops at the right boundary. Include ordinary work, failures you have already seen and cases the team has not tuned against. When a new model, prompt, tool or workflow looks better, run it against the same cases. Check where it improved, where it regressed and whether the gain survives actual use. Google researchers have described tests and monitoring as part of production ML readiness; NIST's AI Risk Management Framework likewise treats governing, mapping, measuring and managing as an ongoing cycle. Sources: Google Research[3] and NIST[10].
That creates a self-improving loop with a human decision inside it: collect feedback, diagnose the cause, change one part, evaluate the candidate, release it carefully and feed the next results back in. The system does not improve because it generated another answer or rewrote its own prompt. It improves when evidence changes the design and the next round of work gets measurably better.
This is the fifth tenet: make learning part of production. The point of a system is not that it never fails. The point is that its failures become legible enough to improve it.
There is a limit to the factory analogy. Physical parts can be held to tolerances; an AI answer may be useful precisely because the task is open-ended. We cannot inspect every answer against a single perfect specification. We can still specify what the system is for, what it may use, what it must never do, when it should ask for help, and what evidence would make us trust it.
That changes the design question. Instead of asking, “Which model should we put into this workflow?” start with: “What should this workflow reliably accomplish, and how will we know when it did?” Only then choose the model, tools, people and checks that make the whole line work.
Sources & notes
Historical manufacturing examples support an analogy, not measured AI outcomes. Google’s 2017 paper supports production ML testing principles. Tesla’s plan is an ambition; its filing reports its experience. NIST AI RMF 1.0 remains published while revision is in progress.