Talking Avatar — document in, talking-avatar video + audio out

gemini-textkrea-2-largegpt-imagegemini-ttschatterbox-ttsinworld-ttslipsyncffmpeg-audio-mix

Output

gemini-text review + speak-out optimize
gpt-image avatar portrait (1024²)
gemini-tts audio (30+ langs)
sync-lipsync/v3 lip-sync
Talking-avatar MP4 (artifact 1)
Standalone audio MP3 (artifact 2)
Multi-speaker mode (optional)

Time

matches source document, typically 30s–5min

Budget

$0.50–$3 per minute of output

Reliability

3.2

Customize this playbook

Fill in the fields below. Your answers are inserted into the complete recipe when you copy it.

local file

for reference; agent auto-detects

what the avatar speaks

free-form description

male | female | neutral

the voice character's apparent age

measured | conversational | energetic | reverent

1:1 (default) | 9:16 (TikTok / Reels) | 16:9 (cinema)

RUNNER

Create and edit images and video with your agent.

npm install -g @livepeer/runner