Voice & Audio
SaaS, open source and skills for voice & audio work — with the limits of each written down.Before you pick
What decides it
Every engine nails a demo sentence. Paste your own text — names, acronyms, numbers — and count retakes across a full hour; retakes bill as fresh generation.
What to avoid
Free tiers, community voices and open checkpoints all sound shippable. Whether you may ship turns on plan tier, model license and the speaker's written consent.
What is changing
Voice quality is converging, open weights included. Control is shifting from markup and sliders to plain-language direction — you describe the delivery, the model performs it.
SaaS
Hosted products you sign up for.Leading
Established picks, at the top of today's aggregate.
- 1ElevenLabsThe default for the most human-sounding TTS, best voice cloning, 70+ languages and voice agents.
- 2SunoThe quality benchmark for AI music generation (full songs with vocals). Suno v5 leads audio-fidelity ELO.
- 3MurfStudio-style TTS for voiceovers, e-learning and presentations. Reddit's top workhorse pick.
- 4SpeechifyBest for listening — turn articles, docs and books into natural speech across devices.
- 5UdioMusic generation rival to Suno, favored by some for vocal texture and remixing.
Emerging
Newer challengers, surfaced by two or more sources.
Open source
Repositories you run yourself.Leading
- 1Coqui TTS (XTTS)The gold-standard open-source TTS / voice-cloning toolkit (~45.8k GitHub★) — XTTS v2 clones a voice from a 3-second clip across many languages.
- 2Bark (Suno)Suno's transformer TTS (~39.2k GitHub★) — highly expressive speech plus non-speech sounds (laughter, sighs, music). A recognized open leader.
- 3OpenVoice (MyShell)MyShell's instant voice cloning (~37k GitHub★) — strong tone/emotion control and low-latency inference. Widely adopted open cloner.
- 4RVC (Retrieval-based Voice Conversion)The de-facto open voice-conversion tool (~36.6k GitHub★) — best balance of quality, speed and ease for real-time voice changing/cloning.
- 5Fish Speech (fishaudio)Fast, high-quality multilingual TTS / cloning (~31.4k GitHub★) — a newer entry gaining strong traction for its speed-quality balance.
Emerging
- 1andyhuo520/openclaw-assistant-mvpElectron desktop voice assistant demo: hold a key, speak, get spoken replies — the email and calendar answers are still mock data.
- 2code-100-precent/LingEcho-AppSelf-hosted platform for building AI voice agents that take live calls, combining speech recognition, LLMs, voice cloning and a knowledge base.
- 3GravityPoet/ChordVoxDesktop dictation tool for macOS, Windows and Linux: hold a hotkey, speak, and transcribed text is pasted at your cursor.
- 4ServeurpersoCom/omnivoice.cppC++ port of the OmniVoice text-to-speech model, running voice cloning and voice design offline on CPU or GPU across 646 languages.
- 5ServeurpersoCom/qwentts.cppC++ port of Alibaba's Qwen3-TTS, doing offline zero-shot voice cloning and prompt-described voices in 11 languages on CPU or GPU.
Skills
Skills you drop into a coding agent.Leading
- 1ElevenLabs Skills (official)Official ElevenLabs Agent Skills bundle — TTS, music, sound effects, voice changer and isolator. Vendor-official but low adoption (~400★) vs the OSS models.
- 2elevenlabs-voices (robbyczgw-cla)Ready-made voice personas/presets (18 voices) over the ElevenLabs API for quick, consistent TTS from an agent. Thin/single-author.
Emerging