Fish Audio
Fish Audio covers text to speech, voice cloning, transcription, voice design, and real time voice APIs in one platform. Its current S2.1 Pro model supports 83 languages and can handle multiple speakers within the same generated conversation.
The editor gives you more control than a basic text box. Instructions such as `[chuckle]`, `[whisper]`, `[emphasis]`, and `[long pause]` can be written directly into a script. Speaker changes, pacing, emotion, and delivery can all be adjusted before generating the audio. This works well for video voiceovers, audiobooks, character dialogue, podcasts, and conversational agents.
Voice cloning starts with a short recording. Clean audio matters, and Fish Audio recommends using several clips when you want a more reliable result. Cloned voices can be public, unlisted, or private. You should only clone a voice when you have permission from the person being recorded.
The public voice library contains a large selection of community voices, while Voice Design can create a new voice from a written description. Fish Audio also includes speech recognition, audio translation, voice changing, story production, audio separation, and sound effect tools.
Developers can connect through REST APIs, WebSockets, or the Python and JavaScript SDKs. Streaming generation supports low latency playback, and the API covers speech generation, transcription, cloning, and voice agent workflows. Pay as you go API access is available without a subscription or minimum commitment.
You can try the web tools without paying. The free plan is limited to personal use, while commercial projects require a paid plan. Developers can test with the free S2.1 model under fair use limits, though that version does not carry the latency and availability guarantees of the production model.
Generated speech still needs a careful listen before publication. Pronunciation, emotional direction, and cloned voice quality depend on the script and source recording. Long narration and dialogue with several speakers deserve an extra review pass.
Features
- Expressive Text to Speech: Adds emotion, pacing, pauses, emphasis, and delivery instructions directly inside a script.
- Voice Cloning: Creates permitted voice replicas from short audio samples, with public, unlisted, and private visibility controls.
- Multi-Speaker Dialogue: Generates conversations with several speakers in a single project.
- 83 Languages: Produces multilingual speech for international content and localization work.
- Speech to Text: Converts spoken audio into written transcripts through the web app and API.
- Voice Design: Creates a new voice from a written description without requiring a reference recording.
- Real-Time Streaming: Delivers generated audio through WebSockets for voice agents and interactive applications.
- Developer Access: Includes REST APIs plus Python and JavaScript SDKs for speech, transcription, cloning, and agent workflows.
- Additional Audio Tools: Includes audio translation, voice changing, audio separation, story production, and sound effects.
TRY IT