FFmpeg
Command-line suite for decoding, encoding, transcoding, filtering and streaming audio and video, supporting a very wide range of codecs and container formats.
A complete multimedia content production environment designed for creators who need to process video, audio, and images at scale. FFmpeg handles all video and audio processing tasks from format conversion to complex filtering and editing operations, while Whisper provides accurate speech-to-text transcription for creating captions, searchable transcripts, and content analysis. Piper offers high-quality text-to-speech for generating narration and voiceovers in multiple languages and voices. n8n orchestrates automated workflows that can process batches of content, generate thumbnails, create social media variants, and manage publishing pipelines. MinIO provides scalable object storage for organizing raw footage, processed outputs, and media assets with S3-compatible APIs that integrate with external services. This bento box solves the complex challenge of managing large-scale content production workflows while maintaining quality and efficiency throughout the creative process.
Command-line suite for decoding, encoding, transcoding, filtering and streaming audio and video, supporting a very wide range of codecs and container formats.
Speech recognition service wrapping OpenAI's Whisper models, transcribing and translating audio across many languages with timestamps, served over an HTTP endpoint.
Fast neural text-to-speech engine that runs entirely on CPU, producing natural speech from small ONNX voice models at real-time speed on hardware as modest as a single-board computer.
Workflow automation platform with 400+ service integrations, extensible in JavaScript. Runs event-driven and scheduled flows across APIs, databases and model endpoints.
Object storage server with an S3-compatible API for on-premises and edge deployments, providing erasure coding, object versioning, write-once retention locks and bucket replication.