full-logo.svg
AI News & Insights

Virtual Try-On: What the Technology Can and Can't Do Yet

Virtual Try-On: What the Technology Can and Can't Do Yet

Capability questions get answered at the wrong resolution. Somebody asks whether this works, receives a yes or a no, and plans a season on it — when the honest answer is that it works reliably on some of your products, unreliably on others, and the split is predictable before you test anything.*

What follows is a way to predict that split from what you already know about your own range, plus a distinction that matters more than any current benchmark: which limits are properties of the method and which are properties of this moment.

Capability is a surface, not a level

There is no single quality level for this technology. There is a surface that rises and falls depending on what the garment physically is, and two brands running identical tools reach opposite conclusions because their ranges sit on different parts of it.

A brand selling structured cotton shirting will report that it works well. A brand selling sheer layered knitwear will report the opposite, using the same software in the same week. Neither is wrong and neither observation transfers.

This is why borrowed assessments are close to useless here. The question is never whether the technology works; it is where your particular products sit on the surface.

The same effect distorts internal conversations. Two people in one company can disagree completely while both being accurate, because one has been working on outerwear and the other on knitwear. Naming the product group alongside any judgment about capability turns that argument into two compatible observations.

Predict from the product, not the photo

The useful predictors are physical attributes, and they are already recorded in your tech packs — the same documents that line drawings, vectors and tech-sheet drafts can be produced from, so the scoring runs on material a team is already maintaining.

AttributeFavorableDifficult
StructureHolds its shape unassistedDrapes softly, changes with handling
OpacityFully opaqueSheer, mesh, open-weave
SurfaceMatte, visible textureHigh sheen, satin, coated, sequined
LayeringSingle layerMultiple visible layers, linings on show
Detail scaleDetails readable at product-page sizeFine stitching, small hardware, micro-prints
Silhouette changeTarget similar to source garmentFitted replaced by voluminous, or the reverse

Score a sample of your range against these six and you will have a usable forecast without generating a single image. A style scoring favorable on five or six is very likely to produce usable output; one scoring difficult on three or more is unlikely to, whatever the tool.

The exercise takes an afternoon and it removes a quarter of the trial and error. It does not remove testing — several styles can be queued in one run, so a test is cheap enough to keep — it narrows what you test. It also tells you something a test cannot: what proportion of your range sits in each group, which is what actually determines whether the workflow is worth building.

Do the scoring before any vendor conversation rather than after. A brand that arrives knowing sixty percent of its range is favorable is having a different discussion from one asking whether the product is any good, and the specifics also make it obvious very quickly whether the person opposite understands apparel.

Which limits are permanent and which are current

The phrase cannot do yet contains two different things, and conflating them leads teams to wait for improvements that will never arrive.

Current limits move. How well fine texture survives, how convincingly a sheer layer renders, how many attempts a difficult silhouette takes — these have improved and continue to. Capabilities are updated frequently; no fixed cadence is promised, so plan on re-testing rather than on a date.

Method limits do not move. A generated image renders an appearance consistent with a pose; it does not compute how a specific fabric behaves. An output built from a photograph carries what that photograph recorded and constructs the rest. These are consequences of what the approach is, not of how mature it is, and no version will change them.

Sort your own blockers into those two columns. Waiting is reasonable for the first kind and a mistake for the second.

The distinction has an organizational use as well. When someone proposes revisiting this next year, the useful reply names which column their objection sits in. A texture concern is worth revisiting; an expectation that the output will eventually answer fit questions is not, and saying so once prevents the same conversation recurring every planning cycle.

Structure and layering: the hardest axis

If one attribute predicts difficulty better than the others, it is how much the garment's shape depends on being worn.

A structured jacket presents largely the same silhouette on a hanger and on a body. A soft-draping dress does not, which means the system is constructing a shape rather than transferring one. That construction is often plausible and it is a construction.

Layering compounds this. Where a lining, an underlayer, or an open front is visible, the system has to resolve how several pieces sit against each other, and each additional relationship raises the chance of something reading incorrectly. Ranges built on layered styling should expect this to be the persistent cost rather than a temporary one.

Surface: sheer, shine, and fine texture

The second axis is what the material does with light.

Sheer fabrics are hard because the system must decide what shows through, which is information the source photograph rarely contains completely. High-sheen surfaces are hard because highlights carry the lighting of their source and travel with the garment into a scene lit differently. Fine texture is hard for a resolution reason rather than an optical one: below a certain size, detail is approximated rather than reproduced.

None of these rules a garment out, and all of them convert into more attempts and closer checking.On a range dominated by satin or lace, that cost is the whole economics of the decision.

Worth separating the two kinds of difficulty when you score, because they behave differently. Optical difficulty produces output that looks wrong and gets rejected quickly; resolution difficulty produces output that looks fine until someone views it larger. The first wastes attempts, the second reaches a product page.

What the boundary never moves on

Independent of any of the above, and independent of improvement.

The output is a visual asset. It does not predict fit. It does not determine sizing. It does not model how a fabric behaves in motion. It does not forecast returns. Those come from measurements, a graded pattern, a physical sample, and your own returns data.

Logos, printed text and small hardware are frequently reconstructed unless the source resolves them clearly, and reconstructed text is often almost right — which is more dangerous than obviously wrong. Every asset containing them needs comparison against the source, with no exception for volume or deadline.

Building a SKU-level expectation before testing

Turn the attribute scoring into something operational rather than leaving it as an assessment.

Tag each style in your range as favorable, mixed, or difficult, and attach an expectation to each tag: favorable styles go straight into generation batches, mixed styles get a single test before committing, difficult styles are photographed conventionally unless someone has a specific reason. Three tags, applied once at range planning, and the workflow stops being a series of individual judgments made under deadline by whoever happens to be holding the file.

The tagging also gives you a way to notice improvement honestly. When the technology moves, it moves the boundary between mixed and difficult, and a range already tagged will show you that shift as a measurable change rather than an impression. The AI Virtual Try-On module in Lightchain AI (apparel AI) covers the favorable and mixed groups from an existing product photo — see AI Virtual Try-On, and Scale E-commerce for how those assets are used across a catalog. For the difficult group, imagery stays with a camera, while the line-drawing and technical-document side is where the design and production workbench earns its place.

Questions apparel teams ask

Is the technology good enough yet? The question has no general answer, because capability varies by garment attribute rather than sitting at a single level. Score your own range on structure, opacity, surface, layering, detail scale and silhouette change, and the answer becomes specific to you.

Which garments should we not attempt? Anything scoring difficult on three or more attributes, particularly sheer layered styles and high-sheen surfaces with fine detail. Those consume attempts without producing usable output, and additional effort does not change the outcome.

Will these limits improve? Some will. Texture fidelity, sheer rendering and silhouette handling have improved and will continue to. The absence of fit prediction and the construction of unobserved detail are properties of the method rather than of its maturity.

Can we test our way to an answer? Testing tells you about the styles you tested. Attribute scoring tells you about the whole range, including the styles arriving next season, which is the more useful thing to hold.

What if our range is mostly difficult? Then the workflow is probably not the right investment right now, and knowing that after an afternoon of scoring is a good outcome. Retag annually rather than re-testing continuously.

Does a better source photograph overcome a difficult attribute? It helps and it does not overcome. A better photograph raises results within an attribute group; it does not move a sheer layered garment into the favorable group.

What the honest answer sounds like

Not yes or no, but a proportion. Roughly this much of your range will produce usable output reliably, this much will need testing, and this much should keep a camera — and you can establish those three figures from tech pack attributes before generating anything. Separate the limits that will move from the ones that will not, tag your range rather than assessing it style by style, and retag when the boundary shifts. That is a more durable answer than any benchmark, because it is about your products rather than about the technology.

Score twenty styles against the six attributes this week.

**Take a representative sample across your range and mark each one favorable, mixed or difficult on structure, opacity, surface, layering, detail scale and silhouette change. The proportion falling into each group is your real answer, and it will hold for next season's range as well, since the attributes travel with the products rather than with the software. → **AI virtual try-on

About the author

[REPLACE — real person's name, role, relevant experience, LinkedIn URL. No team byline.]