How to Generate an AI-Powered Virtual Outfit-Showcase Video from a Single Model Photo

Even a single model image can quickly generate a natural, realistic virtual outfit‑wearing video.
In this tutorial, we’ll upload a model reference image to PixPix, use an Agent to generate video prompts, and then invoke the Seedance 2.5 model to transform the static model photo into a mirror‑front selfie‑style outfit‑display video.
The generated model will naturally turn in front of a full-length mirror, adjust their hair, and showcase the neckline and skirt hem, allowing viewers to examine the garment’s silhouette and details from multiple angles.
Step 1: Prepare the Model Reference Image

First, prepare a complete, clear image of the model.
The reference image used in the video shows the model’s front, side, and back views, helping the AI recognize:
The model’s face and hairstyle
The front design of the garment
The side silhouette of the garment
The back structure of the garment
Details of the neckline and bow
The proportion between the top and skirt
The color and style of shoes and socks
It is recommended to use a solid‑color background or a white‑background model photo, ensuring the person is fully visible from head to toe. Multi‑angle reference images generally help maintain stable garment structure better than a single frontal view.
Step 2: Upload the Model Image to the AI
Enter PixPix’s infinite canvas and upload your prepared model image to the AI.
After uploading, select Agent mode and enter instructions for generating video prompts.
Figure 1 is my model; please provide a 10‑second virtual fitting video prompt—a video shot from a phone held facing a mirror to showcase the clothing and its details. Use a first‑person perspective where the model holds the phone, with the rear camera aimed at the floor‑length mirror. The phone should appear in the mirror reflection but remain positioned to the side of the face, without obstructing the neckline or clothing. Include varied body movements while keeping the outfit consistent and undistorted, in a 9:16 vertical aspect ratio.
This instruction clearly specifies the video’s duration, shooting method, character movements, and consistency of attire, enabling the Agent to generate a more comprehensive video prompt based on the model reference image.
Step 3: Have the Agent Generate Video Prompts
After submitting the instructions, the Agent analyzes the figure, clothing, and styling in the model image and generates video prompts suitable for a mirror‑front selfie scene.
For this video, the generated prompts mainly include the following elements:
Produce a 10-second, 9:16 vertical-format realistic virtual try-on video. The filming method involves the model holding a smartphone and using its rear camera to capture a full-length floor mirror directly in front. The entire video is shot from the phone’s perspective facing the mirror. The mirror simultaneously shows the model, the phone, and the complete outfit; by moving the phone forward and backward, tilting it up and down, and having the model naturally turn sideways, the overall silhouette and details of the garment are showcased.
No text, no subtitles, no dialogue, no voiceover, no logos, and no visible smartphone interface.
The subject remains stationary.
The model’s outfit must not be altered into a jumpsuit, shorts, a skirt, or a two-piece set.
[Relationship Between Phone and Mirror]
The model holds a plain-colored, unbranded smartphone in their right hand, using the rear camera to film the floor mirror. In the mirror, the phone should appear either at the side of the model’s face or above their shoulder, without obstructing the black bow, neckline, waist, or pants.
The right hand always holds the same phone; the left hand is used to adjust the clothing. Only one model and one phone should appear in the mirror, with correct left-right orientation. The phone must not float, change hands, reappear, or stick to the palm.
[Setting]
A bright, clean, modern fitting room or bedroom, with a full-length floor mirror featuring a light wood‑colored thin frame placed against the wall. Light beige walls, light wooden floors, a soft cream‑colored rug, simple low stools, and a few green plants create an uncluttered backdrop, free of scattered clothes.
Soft natural light streaming through the window, realistic skin tones, and accurately rendered gray‑pink hues on the garments. The footage should have an authentic smartphone video feel, with subtle natural handheld shake and automatic focus adjustments, but no severe shaking.
[Timeline]
0–1.8 seconds: Full-body opening shot facing the mirror
The phone is positioned at a wide angle to capture the full-length mirror, showing the model’s complete reflection from head to toe. The model holds the phone in their right hand at the side of their face, ensuring the neckline remains visible, while their left hand hangs naturally. One foot steps slightly forward, shifting the body’s center of gravity to the opposite side, followed by a gentle shoulder sway. The black halter‑neck bow, gray‑pink chiffon top, waistline, one‑piece shorts, stacked socks, and Mary Jane shoes are clearly displayed.
1.8–3.5 seconds: Close-up shots of the neckline and fabric
Keeping the phone pointed toward the mirror, the model takes a small step closer to the floor mirror while gently raising the phone slightly higher and bringing it nearer to the surface, transitioning the composition from a full-body view to a close-up shot of the chest and head.
With their left hand, the model delicately adjusts the black bow, pinching a black ribbon and then releasing it naturally; afterward, their fingertips lightly caress the fine pleats beneath the neckline. The shot clearly reveals the bow’s shape, the black ribbon tail, tiny dark speckled patterns on the pink fabric, and the sheer texture of the chiffon. Movements should mimic everyday clothing adjustments—avoid stiff, rigid gestures pointing at the garment.
3.5–5.2 seconds: Details of the waist and pants
The phone lens smoothly tilts downward from the chest, focusing on the area between the chest and thighs. The model’s left hand glides along the chiffon top toward the waist, gently pressing and releasing the elastic band to highlight the loose, flowing structure of the upper body and the natural waistline.
Next, the left hand lightly pinches the outermost layer of chiffon ruffles on one pant leg, spreading them outward by a few centimeters before releasing them, allowing the ruffles to fall back naturally. It is essential to clearly show both separate pant legs, avoiding any appearance of a single skirt.
5.2–6.8 seconds: Side view presentation
The phone lens is gently raised again and moved slightly backward, returning to a near‑full‑body composition. The model naturally turns her body to the left at about a 45-degree angle, keeping the phone in her right hand at the side of her face and aimed at the mirror.
The left hand rests lightly on the hip, while one leg extends slightly backward. The relaxed, voluminous upper body, the cinched waist, the side profile of the one‑piece shorts, and the double-layered ruffle edges are all clearly shown. The chiffon, bow ribbons, and long hair sway gently with the turn.
6.8–8.4 seconds: Rear side view
The model continues turning to about a 70-degree angle toward the rear, gently guiding her long hair draped over her back to one side with her left hand. She looks back at the phone screen through the mirror, keeping the phone near her shoulder, with the camera still focused on the mirror.
The bare shoulders and back, the halter‑neck design, the tie at the nape, and the black ribbon trailing down the back are all clearly visible, as well as the back silhouette of the one‑piece shorts. The phone must remain in the right hand, and the figure must never completely turn away from the mirror, losing the shooting angle.
8.4–10 seconds: Returning to the front for the finale
The model naturally turns back to face forward, taking half a step closer to the mirror. Her right hand lowers the phone slightly below the level of her face, revealing her full face and the black bow.
Her left hand first tidies the curls around her ears, then gently smooths the chiffon at her waist. With one foot naturally stepping forward, she looks at the phone screen through the mirror and wears a relaxed smile. The shot concludes with a composition focusing on the area above the knees, leaving the black bow, the dotted texture of the chiffon, the waistline, and the ruffled pant legs clearly visible.
[Action Requirements]
The complete motion sequence is as follows: a full-body standing pose, approaching the mirror, adjusting the bow tie, touching the chiffon pleats, lowering the camera angle, showcasing the waist, gently unfolding the ruffled hem of the pants, returning to a full-body view, turning sideways, parting the hair to reveal the back, returning to the front, and finally smoothing the hair and waist.
The movements should be varied yet remain authentic to the act of trying on clothes in front of a mirror—avoid dancing, spinning rapidly, exaggerated twisting, or repetitive mechanical posing. The phone’s movement and the model’s body gestures must coordinate seamlessly, ensuring that the mirror composition always adheres to realistic shooting logic.
[Audio]
Use light, fresh, and stylish instrumental music, supplemented by subtle sounds such as the friction of holding a phone, the rustle of chiffon fabric, the sound of shoe soles moving, and natural ambient room noises. No lyrics, no dialogue, and no voice-over narration.
[Prohibited Actions]
Do not use third-party cameras to circle around the subject; do not abruptly leave the mirror’s field of view; do not include photographers or additional people in the frame; avoid incorrect multiple reflections, infinite mirrors, dual phones, dual hands, or dual models; ensure the phone does not obscure the bow tie or clothing details; prevent jumpsuits from transforming into skirts; prohibit the left and right pant legs from sticking together; make sure black halter straps, bows, waist elastic bands, dotted textures, and ruffled hems remain visible; prohibit changes in garment color, face-swapping, sudden hairstyle alterations, abnormal finger movements, phones slipping through hands, delayed reflections, or any desynchronization between reflections and real‑time actions; and forbid all text, watermarks, logos, or visible camera interfaces.
Once the prompts are generated, copy the complete text for later use in creating the video.
Step 4: Enter the Infinite Canvas to Generate the Video
Return to PixPix’s infinite canvas and connect the model reference image with the video generation area.
After selecting video mode, complete the following settings:
Upload or select the model reference image.
Choose the Seedance 2.5 model.
Set the aspect ratio to
9:16.Set the resolution to
720P.Set the video duration to
10秒.Paste the video prompt generated by the Agent.
Submit the video generation task.
Once the generation is complete, you will obtain a virtual outfit‑wearing video in a mirror‑selfie style.
Step 5: Check for consistency between clothing and the person
The most important aspect of a virtual outfit‑wearing video is that there should be no noticeable changes in either the clothing or the person.
After generation, carefully verify:
Whether the model’s face remains consistent
Whether the hairstyle and hair color have changed
Whether the color of the pink outfit remains stable
Whether the black bow is intact
Whether the neckline appears distorted
Whether the proportions of the top and skirt are consistent
Whether the skirt hem moves naturally with fabric dynamics
Whether shoes and socks remain unchanged
Whether the phone incorrectly obscures any part of the outfit
Whether the direction of mirror reflections is appropriate
Whether fingers and limbs appear natural
Whether the background exhibits flickering or distortion
If only certain movements show issues, you can adjust the prompt and regenerate without replacing the original model image.
Quickly generate a virtual outfit‑wearing video using “Make the Same”
In addition to generating from scratch on the Infinite Canvas, you can also find ready-made AI video examples on the PixPix homepage at PixPix Home .
Scroll down to the “AI Videos” category on the homepage, locate an appropriate mirror‑selfie or outfit‑display video, then click the corresponding “Make the Same” button .
When using a template from another video, simply replace the original reference model image with your own, and you can reuse the existing motion patterns and camera angles.
This method is suitable for:
Those who don’t know how to write motion prompts
Those who wish to quickly replicate a mirror‑selfie outfit effect
Need to batch replace different models or outfits
Want to reuse established video camera movements and actions
Need to quickly create e-commerce apparel showcase content
How can you improve the stability of virtual try-on videos?
Use multi-angle model images
Having the front, side, and back views simultaneously in the reference image allows AI to better understand the garment’s structure, reducing the likelihood of style changes when the model turns.
Avoid overly complex motions
For self-portraits in front of a mirror, it’s best to use small, natural gestures, such as:
Tidy your hair
Gently touch the collar
Lift the skirt hem
Show off your side profile
Lightly tiptoe
Adjust your standing posture
Face the mirror and hold the pose
Quick spins, jumps, or large arm swings are more likely to distort both the clothing and the figure.
Clearly specify that the outfit must remain unchanged
The prompt should emphasize:
No outfit change
No change in clothing color
No change in clothing pattern
No change in collar design
No change in bow tie shape
No change in the number of skirt layers
No change in shoe and sock styles
These restrictions help the model focus on the person’s movements and fabric dynamics.
Control where the phone is positioned
The phone should appear in the mirror reflection but must not completely obscure the collar or main body of the garment.
You can specify in the prompt that the phone should be placed to the side of the face or only partially cover part of the face, while keeping the shoulders, collar, and chest area fully visible.
Summary
This virtual try-on video production workflow can be summarized as follows:
Prepare multi-angle model images → Upload to PixPix → Use the Agent to generate video prompts → Enter the infinite canvas → Select Seedance 2.5 → Set 9:16、720P and 10秒 → Paste the prompt → Generate the video
If you don’t want to design movements from scratch, you can also go to the AI Video category on the PixPix homepage, find a suitable front‑mirror selfie example, and then use the “Make the Same” feature to replace the model image, quickly generating a new virtual outfit video.

AI Image Tool Built for E-commerce Teams
For new product launches, advertising, and promotional campaigns, use AI to generate product images, scene visuals, ad creatives, and short video assets — making content production faster.