Ollama
Local runtime that downloads, quantises and runs open-weight language models, exposing them over its own HTTP API and an OpenAI-compatible endpoint. Also serves embedding models.
Local runtime that downloads, quantises and runs open-weight language models, exposing them over its own HTTP API and an OpenAI-compatible endpoint. Also serves embedding models.
Self-hosted chat interface for local and remote language models, adding document retrieval, model switching, per-user workspaces, tool calling and an extensible pipeline layer.
Visual builder for language-model applications, wiring retrieval, memory, tools and model nodes on a canvas and exposing the finished flow as an API endpoint or embeddable chat widget.
Gateway fronting more than 100 model providers behind a single OpenAI-compatible API, adding routing rules, fallback chains, retries, per-key spend limits and request logging.
Speech recognition service wrapping OpenAI's Whisper models, transcribing and translating audio across many languages with timestamps, served over an HTTP endpoint.
Fast neural text-to-speech engine that runs entirely on CPU, producing natural speech from small ONNX voice models at real-time speed on hardware as modest as a single-board computer.
In-memory data store used as a cache, message broker and ephemeral database, offering strings, hashes, lists, sets, sorted sets and streams, with optional persistence to disk.
Web server and reverse proxy that obtains and renews TLS certificates automatically, supports HTTP/3, and is configured either by a short text file or entirely over an API.
Voice-centric AI interaction platform optimized for speech-based applications and accessibility use cases. Combines speech-to-text, text-to-speech, and conversation management to create seamless voice interactions. This platter addresses the growing need for voice-enabled AI applications, from accessibility tools to IoT integrations and hands-free computing environments.