Video workspaceYour generated video will appear here

Create motion, not just frames

Choose a mode, add the required media, and start with the low-cost settings while you iterate.

Add one portrait

A single JPG, PNG or WebP up to 8 MB. A clear, front-facing face gives the cleanest result.

Write the script or attach audio

Type up to 10,000 characters and choose a voice, or upload MP3, WAV or FLAC up to 10 MB.

Render at 720p or 1080p

Pick a resolution, open Advanced if you want finer control, and generate the talking clip.

A script and a face are the whole workflow

A talking avatar normally means recording a voice, cutting it to length and then hand-matching mouth shapes. Here the portrait and the words are the only inputs. Choose from thirty voices — each marked male or female — across English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean and Hindi, and the speech is produced alongside the video.

Already have a recording? Attach it. Uploaded audio takes priority over the typed script and sets the length of the finished clip. Without audio, the length comes from the script itself at a steady 150 words per minute, so you can size a clip before you render it: about 75 words is about half a minute. Billing follows that same duration, per second, at the rate shown on the cost bar before you press Generate.

The Advanced panel is there when a take needs steering — a video prompt for how the person should behave on camera, a voice prompt for delivery, a negative prompt with an adjustable strength from 0 to 4, and a seed to repeat a run. For characters that need to move rather than speak, see Motion Transfer or EZ Animate.

Talking Avatar — FAQ

What does the AI Talking Avatar do?

It turns a single portrait image into a video of that character speaking. You either type a script and choose a voice, or upload your own audio track, and the portrait is animated to match what is being said.

Do I have to record my own voice?

No. Typing a script is enough — pick one of the built-in voices and the speech is generated for you. Uploading audio is optional, and it is there for when you already have a recording you want the character to perform.

How many voices and languages are there?

Thirty voices, each labelled male or female, across ten language options: English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean and Hindi.

Can I upload my own audio instead?

Yes — MP3, WAV or FLAC up to 10 MB. When audio is attached it takes priority over the typed script, and its length becomes the length of the finished video.

How long will the video be?

If you upload audio, the video is as long as that audio. If you only type a script, the length is estimated from the word count at 150 words per minute, so a roughly 75-word script produces a roughly 30-second clip.

What image works best, and what can I fine-tune?

One clear portrait as a JPG, PNG or WebP up to 8 MB. Scripts can run to 10,000 characters, output is 720p or 1080p, and the Advanced panel adds a video prompt, a voice prompt, a negative prompt with a strength from 0 to 4, and a switch to disable prompt upsampling.