# How AI virtual try-on puts a garment on your photo

Canonical: http://localhost:3000/guides/how-ai-virtual-try-on-works
Updated: 2026-09-24

> Outfello renders AI virtual try-on photos on request, placing a single piece or a whole outfit onto one full-body reference photo. AI virtual try-on works by giving an image-generation model a photo of a person and an image of a garment. It preserves pose, body and face, but it does not simulate fit or size.

You pick a jacket, press a button, and a few seconds later there is a photo of you wearing it, a jacket you may never have put on. AI virtual try-on looks like a camera trick, but it is closer to a painter working from two references: a photo of you and a picture of the garment. Understanding that makes it much clearer what the result can tell you, and what it cannot.

This guide explains how the image is made, what it keeps from your photo, where it tends to go wrong, and why the reference photo matters more than anything else you control.

## Two inputs, one image model

Modern AI virtual try-on systems are built on image-generation models, the same broad family that turns text prompts into pictures. Instead of starting from a sentence, the model is conditioned on two images:

1. **The person image:** a photo of you, which fixes your pose, body shape, face, skin tone and the background.
2. **The garment image:** a picture of the piece, ideally a clean cut-out, which fixes its shape, colour, fabric and details.

Along the way, the system usually works out some structure from the person image: where your body and limbs are, which area the garment should cover (the torso for a top, the legs for trousers), and which parts of the original clothing need to be replaced. The model then generates a new image in which that area is repainted with the garment, while the rest of the photo is kept as close to the original as possible.

This repainting approach is often described as inpainting: the model fills a masked region of an existing photo rather than inventing a whole new scene. It is also why the edges of the garment area matter. If the mask is too small, bits of your original clothing show through; if it is too large, the model redraws parts of your arms or background that did not need touching.

It is not 3D. There is no body scan, no cloth simulation and no measurement of the garment. The model has learned, from a very large number of photos of clothed people, what a garment of a certain kind usually looks like when worn by a body in a certain pose, and it produces a plausible picture of that.

## What the model preserves

Because the model is anchored to your photo, some things come through reliably:

- **Pose.** Your stance and the position of your arms and legs stay as they were.
- **Body shape and proportions.** The garment is drawn around your silhouette as photographed.
- **Face and hair.** Good systems leave the head largely untouched, since it sits outside the garment area.
- **Colouring and light.** Skin tone, the lighting direction and the background carry over, which is why a try-on can answer questions like "does this green work with my colouring?"

This is the useful core of try-on: seeing a colour, a length or a silhouette on your own body, next to your own face, instead of on a model with a different build.

## Where AI virtual try-on fails

The same approach produces predictable weak spots. Knowing them helps you read a result without over-trusting it.

| Weak spot | What you see | Why it happens |
| --- | --- | --- |
| Fit and size | Every piece looks as if it fits | The model draws a plausible worn garment; it has no measurements to be wrong about |
| Patterns and logos | Lettering changes, prints shift or blur | Details are redrawn, not copied pixel for pixel |
| Hands and feet | Odd fingers, merged shoes, missing toes | Small, complex shapes that image models often struggle with |
| Layering | A jacket over a top merges with it, or the top shows through | The model has to invent how two garments interact |
| Garment length | Hems land slightly higher or lower than in reality | Length is inferred from the garment image and your proportions |
| Unusual cuts | Asymmetric or very structured pieces get simplified | The model falls back towards common shapes |

The most important line is the first one. A try-on image says how a piece is likely to look, not whether your usual size will fit. For fit, the garment's measurements and your own are still the only reliable guide.

## Why one good reference photo matters

The person image is the half of the input you control, and it sets the ceiling for every result. A try-on can only be as clear as the photo it starts from.

The model needs to see your body clearly to know where the garment goes. That points to a few practical rules:

- **Full body, head to feet in frame.** Cropped feet mean the model invents them, and shoes are then drawn onto guessed feet.
- **Facing the camera, arms slightly away from the body.** Crossed arms and turned hips hide the areas the garment has to cover.
- **Fitted, plain clothes.** A baggy hoodie hides your shape, so the model has to guess where your torso is. Plain clothes also leave fewer patterns to leak into the result.
- **Even light and a simple background.** Harsh shadows and clutter give the model more to preserve and more to get wrong.

Using one consistent photo also makes results comparable. If every piece is rendered onto the same image, differences between try-ons come from the garments, not from a change in pose or lighting. [How to take a full-body reference photo for try-on](http://localhost:3000/guides/reference-photo-for-virtual-try-on) walks through setting one up.

The garment image matters too, and it benefits from the same clean cut-out that [AI garment extraction](http://localhost:3000/guides/how-ai-garment-extraction-works) produces: a whole piece on a plain background gives the model the clearest picture of what to draw.

## How it works in Outfello

In Outfello, you add one full-body reference photo of yourself, facing the camera with your head and feet in frame. Any piece in the wardrobe can then be rendered onto it on request. Nothing is rendered when you import clothes, so try-ons are only made for pieces you choose.

Each render uses one try-on from your plan's monthly allowance, and if generation fails, the credit is returned. A piece has a limited number of wear-on attempts. In the outfit builder you can also render a whole outfit, jacket to shoes, head to toe onto the reference photo as a single try-on. Allowances differ by plan and are listed on the [pricing page](http://localhost:3000/pricing); imports and try-ons are separate allowances, so running out of one does not block the other. The [virtual try-on](http://localhost:3000/use-cases/virtual-try-on) page covers the feature in more detail.

![A piece open in the Outfello editor, with its modeled photo beside the name, category, colour and tag fields](http://localhost:3000/screenshots/editor.webp "A cut-out beside its wear-on photo")

## Reading a try-on result sensibly

A try-on is best used as a sketch, not a promise. It is good at answering:

- Does this colour work next to my face?
- Is this length right for my proportions?
- Do these pieces look like an outfit together?

It is not good at answering:

- Will this size fit?
- Is the fabric as heavy or as sheer as it looks?
- Will this exact print or logo look like this?

If a result looks wrong in a small detail, such as a hand or a logo, that is usually the model's weak spot rather than a sign the piece will look bad. If it looks wrong across the whole image, the reference photo is the first thing to check.

## Frequently asked questions

### Does AI virtual try-on show whether a garment will fit me?

No. It shows roughly how a piece looks on your body and with your colouring, but it does not measure you or simulate fabric, so it cannot tell you whether a size is right or how tight a waistband will be.

### Why does a logo or print look slightly different in the try-on?

The model redraws the garment onto your photo rather than pasting it, so fine details such as lettering, small prints and exact stripe spacing can drift. Plainer pieces and clear, well-lit garment images reproduce most faithfully.

### What kind of reference photo gives the best results?

One full-body photo, facing the camera, with your head and feet in frame, in even light and fitted, plain clothes. Loose or layered clothing and hidden hands or feet give the model more to guess.

### What happens in Outfello if a try-on fails?

Each render uses one try-on from the plan's monthly allowance, and if generation fails the credit is returned. Nothing is rendered automatically at import, so try-ons are only spent on pieces you ask for.
