Talking Avatar — document in, talking-avatar video + audio out
gemini-textkrea-2-largegpt-imagegemini-ttschatterbox-ttsinworld-ttslipsyncffmpeg-audio-mix
Output
gemini-text review + speak-out optimize
gpt-image avatar portrait (1024²)
gemini-tts audio (30+ langs)
sync-lipsync/v3 lip-sync
Talking-avatar MP4 (artifact 1)
Standalone audio MP3 (artifact 2)
Multi-speaker mode (optional)
Time
matches source document, typically 30s–5min
Budget
$0.50–$3 per minute of output
Reliability
3.2
Customize this playbook
Fill in the fields below. Your answers are inserted into the complete recipe when you copy it.
local file
for reference; agent auto-detects
what the avatar speaks
free-form description
male | female | neutral
the voice character's apparent age
measured | conversational | energetic | reverent
1:1 (default) | 9:16 (TikTok / Reels) | 16:9 (cinema)
RUNNER
Create and edit images and video with your agent.
npm install -g @livepeer/runner