Media to Script — paste image / video / audio, get a human-readable narrative
gemini-textmarlin-videowizpergemini-ttschatterbox-tts
Output
gemini-text (vision) for images
marlin-video for short videos
wizper for audio transcription
gemini-text narrative rewrite
6 reading levels, 5 output formats
Chains into talking-avatar.md
Time
matches source (single-pass, finite media)
Budget
$0.01–$0.50 per source, plus optional voice render
Reliability
4.0
Customize this playbook
Fill in the fields below. Your answers are inserted into the complete recipe when you copy it.
or any video URL; yt-dlp handles fetch
uploaded file
any gemini-text supported language
5yo | 10yo | teen | general-adult | academic | museum-guide
"30 seconds" | "1 minute" | "5 minutes" | "as long as needed"
warm-curious | reverent | clinical | playful | journalistic
script (default) | bulleted-notes | dialogue (multi-speaker) | poem
RUNNER
Create and edit images and video with your agent.
npm install -g @livepeer/runner