Media to Script — paste image / video / audio, get a human-readable narrative

gemini-textmarlin-videowizpergemini-ttschatterbox-tts

Output

gemini-text (vision) for images
marlin-video for short videos
wizper for audio transcription
gemini-text narrative rewrite
6 reading levels, 5 output formats
Chains into talking-avatar.md

Time

matches source (single-pass, finite media)

Budget

$0.01–$0.50 per source, plus optional voice render

Reliability

4.0

Customize this playbook

Fill in the fields below. Your answers are inserted into the complete recipe when you copy it.

or any video URL; yt-dlp handles fetch

uploaded file

any gemini-text supported language

5yo | 10yo | teen | general-adult | academic | museum-guide

"30 seconds" | "1 minute" | "5 minutes" | "as long as needed"

warm-curious | reverent | clinical | playful | journalistic

script (default) | bulleted-notes | dialogue (multi-speaker) | poem

RUNNER

Create and edit images and video with your agent.

npm install -g @livepeer/runner