Motion in generated video is inferred rather than observed. What follows from that depends on knowing what motion is inferred from — and for cloth, the answer is two properties, neither of which appears in a photograph.
Two properties decide how cloth moves
Mass per unit area. How hard gravity pulls on a given piece of fabric. This sets the scale of the movement — how far a hem travels when it swings, how much force it takes to lift, how decisively it falls.
Bending stiffness. How tightly the material can fold. This sets the shape of the movement — whether the fabric forms a few large curves or many small ones, and whether it breaks into creases or rolls.
Together these determine a characteristic rhythm. Heavy and fluid material swings slowly through a wide arc. Light and stiff material flutters quickly through a narrow one. The two properties are independent, so all four combinations exist and each has its own signature.
Everything a viewer recognizes as “that moves like silk” or “that moves like denim” is a reading of these two properties in combination.
That reading happens without training. People who have never thought about cloth mechanics can tell instantly when a garment moves wrongly, because everyone has handled thousands of fabrics and built an intuition for what follows from what. The intuition is unavailable as a description and completely reliable as a detector, which is an awkward combination for anyone trying to specify what they want.
Appearance carries neither
Here is the difficulty for anything working from images.
Two fabrics can look nearly identical and differ substantially in weight. A heavy polyester made to imitate silk photographs like silk and moves nothing like it. A fine wool and a fine polyester can be indistinguishable in a still frame and behave differently the moment either one is disturbed.
Stiffness is equally invisible. Surface, sheen, and weave pattern are what a photograph records, and none of them determine how tightly the material folds — finishing and construction do, and finishing is not something a still image reports.
Which means a model trained on appearance has no path to either variable. It has learned what materials look like, in enormous detail, and the properties governing how they move were never present in what it learned from. This is the general principle that cloth is a structure rather than a surface, appearing in the place where the gap is most consequential.
Four quadrants
Splitting each property into two produces a workable classification, and it does not correspond to how fabrics look.
Fluid (folds tightly) | Stiff (folds loosely) | |
Heavy | Satin, heavy jersey, crepe — slow, wide swings that settle firmly | Denim, canvas, coating — moves as a mass, holds its shape, barely oscillates |
Light | Chiffon, fine georgette — fast, complex, many small folds, floats | Organza, taffeta — springs and rustles, holds volume, returns quickly |
Fabrics in the same quadrant move similarly regardless of what they are made of or what they look like. Fabrics in different quadrants move differently even when they look alike, which is the entire problem restated.
The quadrant is also the useful unit for judging generated output. Asking whether a clip moves like the specific fabric is difficult; asking whether it moves like something in the right quadrant is answerable by anyone who has handled cloth.
The most common generated error is a quadrant error rather than a subtle one within a quadrant. Heavy fabrics come out too light — they float when they should fall — which is consistent with the general tendency of generated output toward the more attractive version, since floating reads as elegant and falling reads as ordinary.
The tell is timing, not shape
Generated fabric usually produces plausible shapes. The folds are the right size, they appear in sensible places, and any single frame looks like a photograph of cloth.
What goes wrong is the rhythm. The motion is too fast or too slow for the material shown, the hem travels too far or not far enough, and the relationship between the disturbance and the response is off.
Viewers detect this without being able to name it. They report that a clip looks strange, or artificial, or that something about it is not right, and they point at the fabric because the fabric is what looks wrong — while the frame they are pointing at is fine.
Which produces a specific reviewing error. Reviewers pause the video and examine the frame, because examining a frame is what they know how to do. The frame is the part the model got right. The failure is in the sequence, and pausing removes exactly the dimension that contains it.
Damping: how it stops
The most reliable single thing to watch is not the movement but its ending.
Real fabric loses energy. It swings, swings less, swings less again, and stops — losing energy to internal friction and to air, at a rate determined by the same two properties. Heavy fluid material settles firmly after a couple of oscillations. Light stiff material returns quickly and rustles. Every material has a settling behavior as characteristic as its drape.
Generated motion fails in two directions here. It stops abruptly, with no return swing, which reads as stiffness the material does not have. Or it continues oscillating until the clip ends, which reads as something closer to liquid than to cloth.
Both are visible in a few seconds if you watch for the ending rather than the movement. Neither is visible in a still.
What a reference clip settles
The way to close this gap is not a better prompt. It is a short clip of the actual fabric moving.
A few seconds of the real material being lifted and released contains both properties directly — the scale of the swing and the tightness of the folds are right there, along with the settling behavior. It is far more informative than any description, and describing these properties in words is close to impossible without becoming technical.
Supplying that clip as a reference constrains the output toward the right quadrant, though what transfers from a reference is partial and each additional dimension arrives weaker. Motion is one of the harder things to transfer, and a reference improves the odds rather than settling the question.
Capturing these clips is also worth doing for its own sake. A library of short movement references, one per material, is useful for briefing photographers, evaluating output, and settling arguments about what a fabric does.
It is also cheap in a way that most reference material is not. A phone, a hand, and a few seconds per material is the entire production requirement, and the result stays valid for as long as the material is in the range. Nothing else in this article costs so little relative to what it settles.
Where this matters, and where it does not
Not every product needs its fabric to move correctly.
It matters most where movement is the product’s appeal: anything fluid, anything cut on the bias, eveningwear, and outerwear with volume. For those, motion is a large part of what the customer is buying and getting it wrong misrepresents the product.
It matters least for structured garments that barely move — tailoring, denim, anything that holds its shape. A jacket that moves as a mass is easy to get approximately right, because the range of plausible behaviors is narrow.
And it does not matter at all for clips where the fabric is not the subject: context, atmosphere, a garment glimpsed in motion at a distance. Generated motion is entirely serviceable there, and the constraints in this article do not apply.
Knowing which of the three a clip belongs to is the decision worth making before generating anything.
FAQ
Can I describe the fabric’s weight in a prompt?
Words like heavy or flowing shift the output in the right direction and do not specify the two properties in the way a physical measurement would. A reference clip carries more.
Why does slow motion look more convincing?
Slowing footage compresses the timing errors, so a rhythm that is wrong at normal speed becomes less distinguishable. It is a reasonable presentation choice and it does not correct the underlying behavior.
Which fabrics are hardest?
Light and fluid materials, since their motion is complex, fast, and highly characteristic. Viewers also have strong expectations about how chiffon behaves, so errors are noticed readily.
Is this improving with newer models?
Output looks more convincing as models improve, and the underlying gap does not close by itself — the two governing properties are still absent from the training material. Better models produce more plausible motion rather than more accurate motion.
Should I shoot reference clips for every fabric?
For every material where movement is part of the appeal. A library built once serves every product using that material and is quick to capture.
Does this apply to 3D simulation too?
Simulation can compute motion correctly, provided it is given measured material properties. That is a different route with a different requirement, and the requirement is data rather than reference footage.
Where this leaves you
Watch how it stops.
Every discussion of generated cloth focuses on how it looks in a frame, which is the part that is now reliably good. The information about whether the material is behaving lives in the sequence — in how far it travels, how long it takes, and how it comes to rest — and none of that survives being paused. A model that learned how silk looks did not learn how heavy it is, and weight is most of what you are watching.
Watch the last second, not the first
Take any generated clip of fabric in motion and skip to where the movement ends. Does it stop abruptly with no return swing, or keep going until the clip runs out? Real cloth does neither — it oscillates a few times and settles, at a rate that is characteristic of the material. That ending is where generated motion gives itself away, and it takes seconds to check. Then shoot a few seconds of the actual fabric moving, once, and keep it. See how motion is generated from a still image.
Written by