The standard answer to generated imagery is a review step. Someone looks at the output before it ships, and the organization treats that look as the safety mechanism.
It is a real mechanism, and it has a precise range. Knowing the range matters more than strengthening the step, because a review operating outside its range does not fail loudly — it approves.
Review is a comparison device
A reviewer looking at an image is not detecting wrongness in the abstract. They are comparing the image to something, and the something is what does the work.
Every check that survives contact with generated imagery is a check where the reviewer holds a reference that exists independently of the image in front of them. Every check that fails structurally is one where the only available reference is that image, or another image from the same source. That single distinction predicts the whole pattern, in both directions, and it is more useful than any list of tells.
The framing is not obvious, because review feels like perception rather than comparison. A reviewer’s experience is of noticing — the thing looks wrong, and the wrongness seems to arrive directly from the image. That experience is accurate for the categories where the reference has been internalized: someone who has handled the range for years does not consciously compare against a catalog, they simply see that a closure is not one the brand uses. The trouble is that the same sensation attaches to categories where no reference exists at all, and it feels identical from the inside. Confidence is not a signal about whether a comparison was available.
Where it still works
Four categories, and they hold up well:
• Brand and product facts the reviewer already knows. That is not our logo, that closure is not one we use, we do not make this style in that length. The reference is knowledge sitting in someone’s head, acquired before the image existed, and a person who knows the range is an excellent detector of things outside it.
• Anything with a written rule attached. Background specification, crop, aspect, safe margins, whether a required element is present. The reference is the document, and review against a document is reliable in a way that review against taste never is.
• Presence and absence against a checklist. Is there a detail shot of the hardware, is there a back view, does the set contain the frames this category requires. The reference is the list, and a list is the cheapest reference there is.
• Gross physical impossibility. Hands, limb counts, reflections that cannot happen, objects intersecting each other. The reference here is a reviewer’s model of how bodies and objects behave — and this category is the one that is eroding, because it relies on outputs being bad in recognizable ways. It is worth listing honestly and worth not relying on for long.
Three of those four are strong and durable. All three are strong for the same reason: the reference is written down or learned, and it is not the image.
Where it fails structurally
The failures share the opposite property. In each case the reviewer’s only reference is an image, so the comparison confirms rather than tests.
Anything about the garment itself falls here. Whether the drape is right, whether the fabric reads as its actual weight, whether the color is the color — a reviewer settles these by comparing the output to a photograph of the garment, or to the source image the output came from, and a generated image agrees with its source by construction. The agreement is not evidence of anything, and it feels exactly like evidence.
Properties belonging to a set rather than to any single image fall here too. Consistency across a group cannot be found by looking at members of the group one at a time, no matter how carefully each one is examined: where single-image tools break.
And errors running toward the flattering fall here for a third reason, which is not epistemic but institutional. Generated output tends to come back tidier, better proportioned, and more composed than the thing it depicts. A reviewer who stops an image for being too attractive is making an argument nobody wants to receive, so the errors that pass review are reliably the ones that improved the picture.
What makes all three hard to correct is that they generate no signal. A review that stops an image produces a visible event: a rejection, a revision, a conversation. A review that approves a wrong image produces nothing — the file moves on, the meeting ends on time, and the only evidence appears later as a return, a complaint, or a sample that does not match the page, by which point nobody attributes it to an approval. The organization therefore receives steady confirmation that its review is working, assembled entirely from the cases where it had a reference.
Adding reviewers does not help
The usual response to a review that missed something is more review — a second pair of eyes, a sign-off from another function, an escalation tier.
Two reviewers holding the same reference produce agreement, not information. They converge faster than one reviewer would have, and the convergence is read as corroboration when it is nothing more than two people applying one standard to one artifact. Where the standard could not settle the question, neither of them can settle it, and now there are two names on the approval.
It also moves responsibility in the wrong direction. A single named reviewer who is unsure tends to ask. Three reviewers in sequence each have grounds to assume the question belonged to someone else, and the image passes through a structure that looks more rigorous than the one it replaced while carrying less doubt.
Supplying references does
The repairable part is not the reviewer’s attention. It is the absence of anything to compare against, and that is a supply problem with concrete answers.
Put the physical sample on the desk beside the screen, and every question about drape, weight, and color becomes answerable — it moves from the failing column to the working column in one step, because the reference is now independent of the image. Write the specified value into the ticket rather than expecting it to be remembered, and the check becomes a comparison against a document. Show a batch as a batch, on one surface, and set-level properties become visible for the first time. Place the unedited original next to the edit, and the extent of the change stops being something the reviewer has to infer.
Each of those converts a structural failure into an ordinary check. None of them requires a better reviewer.
The obvious objection is cost: there is not a sample on every desk, sets cannot always be assembled in time, and originals are not always kept. All of that is true and none of it changes the conclusion — it only changes what the gate is entitled to claim. A gate without the sample cannot answer garment questions, and the useful response is to say so and move that question to a point in the process where the sample exists, rather than keeping a gate that appears to answer it. Narrowing a gate to what it can compare is not a downgrade. It is the difference between an approval that means something and one that means somebody looked.
Which gives the rule worth taking away: a review step with no reference attached to it is not a weak check. It is not a check. It is a signature.
What to attach to each gate
Reviews are usually described by who does them. A more useful description is what each one is holding.
Gate | Reference it needs | Fails silently without it |
Garment accuracy | The physical sample, or a photograph taken independently of the generated output | Drape, weight, and color read as approved when they were only confirmed against themselves |
Specification conformance | The written spec value, in the ticket | Reviewer recalls the value approximately and the approximate version becomes the standard |
Set consistency | The whole set, displayed together | Every image passes individually and the group is wrong |
Extent of edit | The unedited original, side by side | The size of the change is estimated from the result, which cannot show it |
Brand and range facts | The reviewer’s own knowledge, which is why this gate needs a person who has it | A reviewer without range knowledge performs the same motions and detects nothing |
The table is also a staffing observation. The last row is the one gate that cannot be supplied with a document, and it is the gate most often handed to whoever is available. When gates like these get folded into automated sequences, the reference is the part that goes missing first, because it was never written into the step: where verification sits in an agent workflow.
Questions teams ask about image review
Is human review still worth doing on generated images? Yes, on the checks where a reference exists, and it is very good at those. The error is treating it as a general-purpose net rather than as a set of specific comparisons. A review scoped to what it can actually compare is more valuable than a broad one, because its approvals mean something.
Should we add a second reviewer? Only if the second reviewer holds a different reference from the first. Two people with the same reference reach agreement sooner and learn nothing new, while the second name on the approval makes the result look better supported than it is. If the second reviewer brings the physical sample, that is not a second reviewer — that is a reference, and it would work with the first reviewer alone.
What is the cheapest reference to add? The unedited original, placed beside the edited version. It costs nothing, it already exists, and it converts a guess about how much changed into an observation. Most review setups do not show it, because the output arrived on its own.
Can review catch a problem that affects a whole batch? Not by examining images one at a time, which is how nearly all review is organized. Batch problems are properties of the group and become visible only when the group is displayed as a group. A per-image process can pass every member of a batch that is wrong as a batch.
Why do reviewers approve images that are wrong about the garment? Because they are comparing the image to another image rather than to the garment. The generated result and its source agree by construction, so the comparison returns a match, and a match feels like confirmation. Nothing about the reviewer’s diligence changes this; the reference is the problem.
Can the same review step cover the copy as well? It should not, because the failure works differently there — text can imitate the register of a verified statement perfectly, so reading it carefully does not separate the checked claims from the invented ones: what generated copy can and cannot claim. Treat copy as its own gate with its own reference, which is the list of facts someone confirmed.
Where this leaves you
Human review works exactly as far as the reviewer’s reference reaches, and not one step further. Three of its four strong categories hold because the reference is written down or learned rather than looked at, and every structural failure traces to a check where the only thing available to compare against was another image. The improvement worth making is not a stricter reviewer or a longer approval chain. Go through the gates you already run, ask what each one is holding, and supply the ones holding nothing.
Ask what each review gate is holding
Take your current approval chain and write down, for every gate, the reference the reviewer compares against. Some will have a document, a checklist, or a sample. Some will have nothing, and those are the ones producing signatures rather than checks. Fix them by supply rather than by escalation: the sample on the desk, the spec value in the ticket, the set shown as a set, the original beside the edit. Then confirm that automated steps carry their reference with them.
Written by