Fish Speech (fishaudio)
Leading#5 in Open sourcehigh confidence
Fast, high-quality multilingual TTS / cloning (~31.4k GitHub★) — a newer entry gaining strong traction for its speed-quality balance.Our read
Why it ranks #5
Around 31k GitHub stars and rapid adoption in local-TTS communities place it among the leading newer multilingual speech models, with notable recognition for its speed-to-quality balance.Here is the catch
hardware requirements remain substantial for best performance
installation and checkpoint choices can confuse new users
model licensing is more restrictive than permissive OSS code alone suggests
Does this well
strong naturalness and multilingual cloning quality
efficient inference relative to many large speech models
active development and a growing ecosystem
Pricing
Checked by hand on 2026-07-23. Prices in this category change often — if this looks wrong, it probably is.
Key features
multilingual text-to-speechzero-shot voice cloninglong-form speech generationfast local inference
Quick facts
More in this area
The rest of the Open source column.- 1Coqui TTS (XTTS)The gold-standard open-source TTS / voice-cloning toolkit (~45.8k GitHub★) — XTTS v2 clones a voice from a 3-second clip across many languages.
- 2Bark (Suno)Suno's transformer TTS (~39.2k GitHub★) — highly expressive speech plus non-speech sounds (laughter, sighs, music). A recognized open leader.
- 3OpenVoice (MyShell)MyShell's instant voice cloning (~37k GitHub★) — strong tone/emotion control and low-latency inference. Widely adopted open cloner.
- 4RVC (Retrieval-based Voice Conversion)The de-facto open voice-conversion tool (~36.6k GitHub★) — best balance of quality, speed and ease for real-time voice changing/cloning.
- 6Chatterbox (Resemble AI)Resemble AI's open TTS (~25.7k GitHub★) — a top 2026 pick for expressive, production-grade local speech with emotion control.
- 7F5-TTSA fast diffusion-transformer TTS with strong zero-shot cloning (~15k GitHub★) — notable and actively used, below the top repos on adoption.