text-to-speech

12 servers · 112★ total

Open-source skills that empower any AI agent (Claude, Cursor, Hermes, etc.) to generate end-to-end marketing campaigns - UGC videos, ad videos, product photography and more.

Let your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.

🦞 Headless real-time Python voice assistant powered by ElevenLabs, Gemini, Mem0, MCP tools, and local wake-word detection.

MCP server for Kokoro text-to-speech, with adjustable voices/speed and an optional OpenAI-compatible (kokoro-fastapi) backend.

Speekify turns CLI text, stdin, local .txt/.md/.pdf files and YouTube video transcripts or the readable content of a URL into a local WAV file

OpenWand - A hotkey-driven AI overlay for your desktop. Press a key, pick an intent, and OpenWand reads the right context, then streams an answer without making you leave what you're doing. Local-first, voice in/out, bring your own model provider.

Python

MCP server for text-to-speech generation with Gemini TTS

Stateless NPC intelligence with layered memory cycles, personality evolution, voice interaction, and MCP-based agency in game environments.

MCP server for VOICEPEAK text-to-speech synthesis

Self-hosted operator console for Magnific / Freepik AI. REST + MCP, 48 image models, 52 video models, credit-safe job engine

Self-hosted speech server in one Docker image. OpenAI-compatible /v1/audio/transcriptions and /v1/audio/speech across 12 ASR models (Whisper, Parakeet, Canary, Sherpa-ONNX, Vosk) and 3 TTS engines (Kokoro, Qwen3-TTS voice cloning, Chatterbox Turbo). Live WebSocket ASR, file staging, MCP built in. CPU + CUDA images.

Local AI asset generation MCP server: end-to-end text-to-image/audio/speech and image/text-to-3D with Qwen3-TTS, Stable Audio Open, SSD-1B, and TripoSR.