Talking Head
Generate a lip-synced talking-head clip from a portrait and a voice track — face and voice stay matched.
Scenarios
Tags
Estimated duration
Usually 5-12 min
30-day retention before cleanup
Results, submitted inputs, and uploaded file metadata are kept for 30 days after execution, then cleaned up.
HistoryGenerate one 5-second talking-head clip from a portrait photo and a voice track: the face matches the portrait, the speech matches the audio, with optional subtitles and an emotion direction. Longer voiceovers are made chunk by chunk — run each chunk separately and chain them with the previous clip's last frame.
-
Prepare: a clear portrait (required) and the voice track (required, MP3/WAV/M4A, 5 seconds or less per run). Optional emotion direction and subtitle toggle.
-
You get: the clip in a gallery with player and download, plus a creation note. Generation usually takes 5-10 minutes.
-
Use it for: spokesperson clips, podcast teasers, and product voiceovers.
Output sample
Talking-head preview What the result contains - 5s lip-synced clip with player and download - Portrait and voice matched - The exact prompt used - Emotion and subtitles applied The final report adds evidence, reasoning, and recommended next actions.