How to make an AI UGC video ad

Make an AI UGC video ad from scratch: an AI creator holding your product and talking to camera, built out of short clips so the same face and the same voice carry the whole ad. No actor, no studio.

Guide

This is how to make an AI UGC video ad from scratch: a real-feeling, phone-shot video ad where an AI creator holds up your product and talks to camera, with the same face and the same voice from the first clip to the last. You borrow the shape of an ad that already works, write a fresh script for your own product, and build it out of short clips using two AI models. You come out with a reusable cast folder and a first cut you could post.

The tools

  • ChatGPT Images 2.0 to make and tidy your creator and product images.
  • Seedance 2.0 to turn a still into a talking clip and hold the voice steady.
  • Higgsfield, which gives you both models in one place, with saved characters and props.
  • CapCut to join the clips.
  • Captions (the phone app) for subtitles at the end.
  • Meta Ad Library to study an ad structure that is already earning before you write.

Versions move fast. If a model has a new name by the time you read this, the four moves hold: still image, short clip, one saved voice, edit.

The two models

ChatGPT Images 2.0

Builds and cleans up your creator and product stills, so you begin from a believable image rather than a rough screenshot.

Seedance 2.0

Brings that still to life as a talking clip and keeps the same voice, movement and look across every shot. Both models live inside Higgsfield, so you work in one place.

Four honest rules before you start

  • Write your own script. Learning from how a good ad is built is fair. Reusing its actual lines is not, so keep the shape and fill it with your words.
  • Only claims you can stand behind. Take every claim from the real product page. Do not invent a result, a review, or a number.
  • Tool names change, the habit does not. Still, short clip, one saved voice file, edit.
  • Label AI people when the platform asks. A few platforms now want a synthetic creator flagged. Check the one you are posting to.
Phase 1

Reference ad and script

Find an ad that works. In the Meta Ad Library, look for an ad in your category that has been running for a while, which usually means it is paying its way. You are learning from its shape, so what you borrow is the structure and the words stay yours.

Get a transcript with timestamps. Run that ad through a transcription tool so you can read each line and see how long it lasts. The easiest ones to learn from are plain and personal: a single person to camera, a problem, one bit of proof, a simple ask.

Write your own version. This is the part it is tempting to skip. Keep the running order and write every line yourself for your own product. A solid order to follow: open on the hook, say what it is, show it in use, name the problem it solves, give one honest proof, then land a single ask. Take the proof from your product page and nowhere else.

Split the script into single shots. One shot is one moment. Seedance makes clips of roughly five to fifteen seconds, and a person speaks at about three words a second, so each shot lands around eight seconds and a couple of dozen words. A longer line becomes two clips you join later. Mark each line as said to camera or spoken over pictures, since that sets how you prompt it.

Phase 2

Build your creator and product

Picture your actual buyer. Work out who really buys this and let the creator look like that person. Someone the viewer recognises as a peer will land better than a glossy influencer, so match the age, the room and the ordinary feel to your buyer.

Make the creator in ChatGPT Images 2.0. Write a detailed prompt that asks for realism: real skin, real hair, a plain room, phone-camera framing. Make a few and keep the one that looks least AI-generated. Here is the shape, written for a made-up calm-morning tea brand.

PromptCreator image (ChatGPT Images 2.0)
Vertical phone video still, 9:16, framed head and shoulders at eye level.

An ordinary woman in her mid-forties at home, standing in her kitchen with the phone held out in front of her, mid-sentence, as if she is filming a quick review for a friend.

Make her look real rather than polished: everyday skin with a bit of texture, a few soft lines around the eyes, little or no makeup, hair slightly undone. She has on a comfy jumper.

Behind her the kitchen is bright but blurred, a kettle on the side and daylight through a window. It should read like a clip caught on a normal front camera: a touch soft, a little grain, no studio lighting and no retouching.

Save the creator as an Element. In Higgsfield, open Seedance, add the creator image as an Element, name it, and set its type to Character. After that you bring her into any clip by her tag instead of describing her again.

Add the product as its own Elements. Take clean product photos from the brand site and save each as an Element as well. You usually want a few: the front of the pack for hero shots, the back for the ingredients moment, the product in use for the how-it-works shot. Tag the right one in each clip.

Make separate stills for new poses or rooms. A talking-head look, a mirror look and a sofa look are three starting images. Trying to make Seedance bend one still into a new body position tends to go wrong, so build the right still first and animate that.

Phase 3

Generate every clip

Keep prompts short. Seedance handles a plain prompt better than a crowded one: one action, one accent, the tagged elements, the line. Your creator's first clip has no saved voice yet, so this is where her voice is set.

Said to camera
PromptClip said to camera
@YourCreator looking at the camera, holding @YourTea.
British English accent.
Dialogue: "I used to open five apps before breakfast. Now I make one cup of this and sit for a minute."
Add nothing on screen. No subtitles, no text, no music.

Tell it which kind of line it is. Put the label in front of the words. Dialogue means she is on camera and lip-synced. Voiceover means the words play over a shot, like a product close-up, with the mouth left alone.

Spoken over pictures
PromptClip over pictures
Close-up, filmed like a phone selfie: a hand lifts the @YourTea tin off a bright kitchen counter.
British English accent.
Voiceover: "If your morning already feels loud, you do not need a twelfth app."
Add nothing on screen. No subtitles, no text, no music.

Close every prompt by asking for a clean frame. Tell it to add no on-screen text and no music. That way each clip comes out bare and you put subtitles on once, properly, in CapCut, with no baked-in text or stray music to fight.

Set the voice on the first clip, then reuse it. When your opening clip sounds right, download it, drop it into CapCut, and export the audio on its own as an MP3. Keep it in your project folder. On every clip after that, upload the same MP3 in Higgsfield under Uploads, and Seedance holds the same voice, tone and accent through the whole ad, even on another day.

PromptClip with the saved voice file
@YourCreator looking into the camera.
British English accent.
Dialogue: "This is the only tin on my desk now."
Add nothing on screen. No subtitles, no text, no music.

Attach: your-creator.png as the reference, your-creator-voice.mp3 as the voice.
Aim for about 8 seconds, sized to the number of words.

Keep one moment to one clip. Size each clip to its speech. If a line runs past about ten seconds, make it as two clips and join them in CapCut instead of forcing it into one.

Hold the room steady with a start frame. If she should stay in the same kitchen across clips, take a screenshot of a finished clip in that room, upload it, and mark it as the clip's Start Frame. Every clip from that frame keeps the same background and light.

Give one instruction per line. Plain gestures work best: hold the tin up to camera, turn it to show the front, then look back and talk. Ask a clip to do two things at once and it usually muddles both.

Phase 4

Edit in CapCut

  • Line them up and trim the silence. These clips nearly always open and close on a beat of quiet. Zoom into the audio and cut those gaps. This one step does the most for the pace.
  • Lift the speed a little. The output tends to run slow, so set each clip to about 1.1 times, or 1.2 for anything still dragging, and it starts to sound natural.
  • Add overlays and a bit of pop. Drop a before still onto the canvas and place it by hand. Use a quick transition into the second clip. For emphasis, take a half-second, blow it up to around 130 percent, and let it snap back.
  • Export to post. Vertical, at a size the platform likes. Leave CapCut's quality on its defaults.
Phase 5

Add captions

Send the finished edit to your phone and open it in the Captions app. Choose a clean preset with a solid background, size the text so it reads on a small screen, keep it off the mouth, and export. Since every clip came out with no on-screen text, the subtitles go on in one clean pass. CapCut has its own auto-captions if you would rather stay in one app.

The habit that pays off next time

Hold on to the cast folder. Your next ad is a new script with the same creator image and the same voice file, so the second one takes a fraction of the time the first did. You end up with a creator you can keep filming, ready for the next product.

Master checklist

  • Found a working ad in the Meta Ad Library and pulled a timestamped transcript.
  • Wrote your own script to the same running order, every claim taken from the real product page.
  • Split the script into single shots of about eight seconds, longer lines marked to stitch.
  • Pictured the real buyer and made the creator in ChatGPT Images 2.0, keeping the least AI-looking take.
  • Saved the creator as a Character Element and each product photo as its own Element.
  • Made the first clip with no voice file, then exported its audio as the creator's voice MP3.
  • Made every later clip with the voice file, the reference image and the tags.
  • Marked lines to camera as dialogue and lines over pictures as voiceover.
  • Closed every prompt by asking for no on-screen text and no music.
  • Set a Start Frame for any room that repeats across clips.
  • Edited in CapCut: silence trimmed, speed lifted, exported to post.
  • Captioned in the Captions app and ran the honesty check: true claims, AI labelled where asked, your own words.

FAQ

Why does my creator sound different in every clip?

Because the voice was never locked. Make one clip first, export its audio as an MP3, then upload that same file with every later clip. Seedance uses it to hold the same voice, so the creator stays the same person across the whole ad.

Can I make the whole ad in one go?

Keep each clip to a short bite of speech, roughly fifteen seconds or less, then edit them together. Long single takes drift on the face and the voice, and short clips are far easier to redo when one comes out wrong.

Do I need special gear or a paid studio?

No. The two models, a free editor and a captions app cover the whole job. The work is in the writing and the editing. The kit barely matters.