Batch-generate images via OpenAI Images API using a random prompt sampler and an `index.html` gallery for easy viewing.
Media & Creative
Images, audio, video processing, and creative tools
10 published skills sit in Media & Creative. Every listing shows what the skill does, which runtimes it supports, and how it runs.
- Published skills
- 10
- Runtimes covered
- 9
- Distinct tags
- 28
Narrow these skills down
Filter by runtime
How these skills run
Media & Creative skills (10)
You receive fully synchronized, SEO-optimized show notes, detailed timestamps, and accurate transcripts for your podcast episodes.
Generate spectrograms and feature-panel visualizations from audio using the songsee CLI tool.
Local text-to-speech conversion using sherpa-onnx, enabling offline and private TTS functionality without relying on cloud services.
Transcribe audio to text locally using the OpenAI Whisper CLI tool, without requiring an API key.
ElevenLabs text-to-speech with a Mac-style 'say' UX, enabling quick audio generation and playback from the command line.
Transcribe audio files using OpenAI's Whisper API via a simple bash script, requiring an OpenAI API key.
Search GIF providers, browse in TUI, download, and extract stills/sheets for quick review and sharing.
Control Sonos speakers on your local network for playback, volume, grouping, and more.
Generate and edit images using Gemini 3 Pro Image (Nano Banana Pro) with text prompts and optional image inputs.