Audio Generation
AI speech synthesis and music generation tools
ElevenLabs
ElevenLabs
State-of-the-art AI voice synthesis and voice cloning with support for many languages and emotional expression.
Suno
Suno
An AI music generator that creates complete songs from a text description.
Udio
Udio
An AI music generation platform known for its high-quality output.
Descript
Descript
AI-powered audio and video editor where you edit recordings like a document: delete words to cut footage, clone voices for corrections, and clean up studio sound automatically.
Murf AI
Murf
AI voiceover studio for professional narration: 120+ realistic voices in 20+ languages, with timing sync to slides and videos, team workspaces and an API.
PlayHT
PlayHT
A text-to-speech platform with ultra-realistic voices, voice cloning and a developer API for real-time speech generation.
Speechify
Speechify
A text-to-speech reader that turns documents, articles and web pages into natural audio, popular for accessibility and speed-listening.
AIVA
AIVA
An AI composer that creates original soundtrack music in dozens of styles, with editable tracks and licenses for commercial use.
Soundraw
Soundraw
A royalty-free music generator that composes customizable tracks by mood, genre and length, licensed for videos and podcasts.
Boomy
Boomy
An AI music creator that generates complete songs in seconds and can distribute them to streaming platforms with royalty sharing.
Beatoven.ai
Beatoven
An AI soundtrack generator for video and podcast creators, composing mood-matched, royalty-free background music scene by scene.
Podcastle
Podcastle
An AI-powered podcast studio: studio-quality voice enhancement, text-based audio editing, remote interviews and AI voices in one workspace.
Resemble AI
Resemble
An enterprise voice-cloning platform for games, film and applications, with emotion control, real-time speech APIs and deepfake detection tools.
Listnr
Listnr
A text-to-speech and voiceover platform with 900+ voices, podcast hosting and embeddable audio players for blogs and sites.
WellSaid Labs
WellSaid Labs
An enterprise text-to-speech platform producing broadcast-quality AI voices for corporate training, advertising and products, with a focus on consent-based voice models.
Krisp
Krisp
An AI noise-cancellation layer for calls and recordings, removing background voices and noise in real time plus echo and accent smoothing.
Voicemod
Voicemod
A real-time AI voice changer for gaming and streaming, with hundreds of voice effects, a soundboard and text-to-song generation.
VoiceMaker
VoiceMaker
A text-to-speech platform with hundreds of voices and neural effects, built for converting articles and scripts into downloadable audio files.
Cleanvoice AI
Cleanvoice
An AI audio cleaner for podcasts: removes filler words, mouth noises, dead air and stuttering from recordings automatically.
Dubverse
Dubverse
An AI dubbing platform that re-voices videos across dozens of Indian and global languages, with subtitle generation included.
ElevenLabs v3
ElevenLabs
The newest ElevenLabs model with audio tags for emotion control — laugh, whisper, pause — plus 70+ languages and dialogue generation in a single pass.
ElevenLabs Music
ElevenLabs
ElevenLabs' entry into AI music: generate studio-quality songs from text prompts with full control over style, lyrics and structure.
Sonauto
Sonauto
An AI music generation platform that creates full songs with vocals from prompts and lyrics, competing on quality and creative control.
Kits
Kits AI
An AI music toolkit for artists: train custom voice models, generate AI vocals and harmonies, convert stems into instrument performances and master tracks in the browser.
Stable Audio
Stability AI
Stability AI's music and sound-effects generator: text-to-audio up to full-length tracks with commercial licensing on paid tiers and an API for game and app integration.
Whisper
OpenAI
OpenAI's open-source speech recognition model: multilingual transcription, translation and timestamps that runs locally — the quiet engine inside half the transcription tools on the market.
Cartesia
Cartesia
Real-time voice AI infrastructure: the Sonic TTS family generates natural speech in as little as ~40ms — low enough for phone calls and live agents — alongside voice-changing and cloning APIs.
Moises
Moises (Music.ai)
The musician's practice AI: split any song into stems (vocals, drums, bass), change pitch and tempo without artifacts, detect chords, and practice along — used by millions of players and producers.
Riverside
Riverside
The remote recording studio: captures lossless local audio and 4K video for podcasts and interviews, then AI Magic Clips turns the session into social-ready shorts and show notes.
AssemblyAI
AssemblyAI
Speech AI as an API platform: production-grade transcription plus audio intelligence — speaker labels, sentiment, topic detection, summarization — powering call centers and media pipelines.
Deepgram
Deepgram
Speech-to-text and voice-agent APIs built for speed and scale: end-to-end deep learning models that transcribe in real time at a fraction of legacy costs, powering call centers and live voice apps.