Add one portrait
A single JPG, PNG or WebP up to 8 MB. A clear, front-facing face gives the cleanest result.
Choose a mode, add the required media, and start with the low-cost settings while you iterate.
Upload one portrait, type what it should say, and pick a voice. Your character delivers the line on camera — no recording session, no lip-sync work, no rig.
A single JPG, PNG or WebP up to 8 MB. A clear, front-facing face gives the cleanest result.
Type up to 10,000 characters and choose a voice, or upload MP3, WAV or FLAC up to 10 MB.
Pick a resolution, open Advanced if you want finer control, and generate the talking clip.
A talking avatar normally means recording a voice, cutting it to length and then hand-matching mouth shapes. Here the portrait and the words are the only inputs. Choose from thirty voices — each marked male or female — across English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean and Hindi, and the speech is produced alongside the video.
Already have a recording? Attach it. Uploaded audio takes priority over the typed script and sets the length of the finished clip. Without audio, the length comes from the script itself at a steady 150 words per minute, so you can size a clip before you render it: about 75 words is about half a minute. Billing follows that same duration, per second, at the rate shown on the cost bar before you press Generate.
The Advanced panel is there when a take needs steering — a video prompt for how the person should behave on camera, a voice prompt for delivery, a negative prompt with an adjustable strength from 0 to 4, and a seed to repeat a run. For characters that need to move rather than speak, see Motion Transfer or EZ Animate.