Automating a design workflow removes the pauses. That is the point of it, and it is also where the difficulty is, because the pauses were doing something nobody had assigned them.
Verification used to be free
In a workflow made of tools, checking happened as a side effect. Work stopped between steps — a file was exported, sent, opened by somebody else, waited on — and at every one of those stops a person looked at what they had received.
Nobody scheduled that. It was not a quality process, and it was not written down anywhere. It happened because handovers involve looking, and looking catches things.
An agent chain removes the handovers. The output of one step becomes the input of the next without anyone opening it, which is the efficiency being purchased. What goes with it is every observation that used to happen for free.
The correct response is not to reinstate the pauses. It is to recognize that verification has changed from a byproduct into a cost, and costs have to be placed deliberately.
This also explains a pattern that confuses people evaluating automation. A team automates a process, quality drops, and nobody can point to a step that got worse — every individual operation is performing as well as it did or better. Nothing got worse. Something that was never counted as part of the process stopped happening.
Two tests for whether a step can go
A verification step can be removed if both of these hold.
The error stays cheap. Whatever this step checks can be checked again later, at a similar cost, if it turns out to be wrong. Anything downstream of cutting fabric, placing an order, or publishing fails this test, because after those points a correction costs materially more than it did before.
The error stays visible. This is not the only stage at which the problem could be detected. If later stages cannot see it, removing this check removes the last opportunity rather than one of several.
The second test is the sharper one, and it is the one people apply least. Cost is intuitive; invisibility is not.
Some errors do not get more expensive — they disappear
An error that becomes expensive downstream is at least still findable. An error that becomes invisible is a different category.
Downstream steps take their input as given. A material specification that was estimated rather than measured arrives at the next step as a material specification, indistinguishable from a measured one. Nothing in it says it was a guess. The step after that treats it as established, and by the time anything physical exists, the assumption has been embedded in decisions that were reasonable given what they were told.
The same applies to a claim in copy that came from outside the brief, to grading intent that was never documented, and to a color that was approved on an uncalibrated screen. In each case the problem is not that the error grows. It is that the marker identifying it as uncertain does not travel with it.
Which produces the rule: check anything whose uncertainty is not recorded in the artifact itself, at the point where it is still known to be uncertain.
Chains compound assumptions
There is a second effect that agent workflows produce and tool workflows do not.
In a chain, each step builds on the last. An unverified input at the first step is used by the second, which produces an output the third depends on, and so on. By the end of the chain, the original assumption has been built upon several times.
What that does is not make the error larger. It makes it look more supported. A result that five subsequent operations were derived from appears well-founded, because everything around it agrees with it — and everything around it agrees with it because it was derived from it.
This is the specific way a long chain can be more dangerous than a long manual process. The manual process has the same dependency structure and also has people who occasionally say something looks off.
The compounding has a practical consequence for how errors get investigated. When something is found wrong at the end of a chain, the instinct is to examine the last step, because that is where the problem appeared. The last step is usually correct — it processed what it was given. Working backwards through a chain is slower and is the only method that finds where an assumption entered.
The steps that cannot be removed
Five, drawn from the failures this series has documented elsewhere. Each is listed with the reason it fails one or both tests.
• Where material data came from — measured or estimated. Fails the visibility test: the distinction is not recorded in the output, and every downstream step treats an estimate as a specification.
• Comparison against something physical — at least once, before anything is committed. Fails the visibility test: outputs derived from one another will agree, so no internal check can substitute.
• Grading intent — whether it exists as a document. Fails both tests: it is not carried by any image or model, and by the time a size run is produced, the error is across a shipment.
• Claims in copy against a source — every statement about material, care, or origin. Fails the visibility test, since generated text states an invention in the same register as a fact.
• The substitution decision — whether a digital artifact is being accepted in place of a physical one, and by whom. Fails the visibility test: if it is not recorded, nobody can later establish what was approved on what basis.
The list is short because most steps genuinely can be automated, and that is worth saying plainly rather than treating automation as something to be minimized. These five share a property: the artifact does not carry the information needed to check them later.
Where to put the fewest checkpoints
Each checkpoint interrupts the automation, which is the objection people raise and it is a fair one. So placement matters more than quantity.
Two placements do most of the work. Immediately before any irreversible step — cutting, ordering, publishing — because that is the last moment corrections are cheap. And immediately after any step that introduces information from outside the chain, since that is where unverified input enters and where its uncertainty is still known.
Everything between those points can usually run uninterrupted. A chain with two well-placed checks is stronger than one with six placed by intuition, and considerably less irritating to work with.
The mistake to avoid is placing checks where they are easy rather than where they are needed. Convenient checkpoints tend to cluster at the start, where the work is cheap to redo and nothing has been committed — which is precisely where a check is least necessary.
What an agent should hand back
The most useful change is not a checkpoint at all. It is requiring the chain to report what it assumed.
An agent workflow that returns only a result has hidden its inputs. One that returns the result plus a short account of what it took as given — which material record, which body reference, which claims came from a source and which were generated — makes verification possible without interrupting anything.
Four things are worth having in that account: what was used, where each input came from, what was assumed where an input was missing, and which outputs depend on which assumptions.
That last item is what makes a discovered error actionable. Knowing that a material record was wrong is useful; knowing which of the chain’s outputs were derived from it tells you what has to be redone.
FAQ
Does this mean long agent chains are a bad idea?
Long chains are efficient and they concentrate the consequences of an early error. The response is to place checks at the points identified above rather than to shorten chains, since a shorter chain with no checks has the same problem over fewer steps.
Can an agent verify its own work?
It can check internal consistency, which is exactly the check that cannot detect the failures described here — outputs derived from a shared assumption will agree with each other. Verification has to reference something the chain did not produce.
Who should own the checkpoints?
Whoever is accountable for the consequence of the error each one catches, which usually means different people at different points. A single reviewer at the end is the arrangement that looks tidiest and catches least.
What if the team resists interruptions?
The objection is usually to badly placed checks rather than to checking. Two checkpoints in the right places are typically accepted where six in convenient places are not.
How do we know if an assumption was made?
Only if the chain reports it. Where it does not, the practical fallback is checking the inputs you know are frequently missing — material data and body references being the common cases.
Does this apply to a single generation step as well?
The two tests apply anywhere. A single step has fewer places for an assumption to hide, so applying them is quicker rather than unnecessary.
Where this leaves you
Ask which of your checks were scheduled and which were happening because somebody had to open a file.
The second group is the one automation removes, and it is usually the larger of the two. Nothing about the risk changed — the same errors are possible, in the same places, with the same consequences. What changed is that the observations catching them were a byproduct of a process that no longer stops, and byproducts do not survive being optimized away.
List the checks nobody scheduled
Walk through your current workflow and mark every point where a person opens something. Not the formal reviews — the moments where a file gets received, unpacked, or passed along. Those are the observations you will lose when the steps are automated, and most of them are not written down anywhere as checks. That list is what has to be replaced deliberately, and it is nearly always longer than the list of checks anyone would have named. See how design agents structure a workflow.
Written by