AI Audio
Speech-to-text, text-to-speech
14 tools in this category
Whisper (OpenAI)
Python • MIT
OpenAI speech recognition
WhisperX
Python • MIT
Fast transcription — batch processing, speaker diarization
Bark (Suno)
Python • MIT
Text-to-audio with voice cloning
Retrieval-based Voice Conversion
Python • MIT
Real-time voice conversion
AudioCraft (Meta)
Python • MIT
Music generation — Meta open-source models for audio creation
F5-TTS
Python • Apache-2.0
Natural text-to-speech — flow matching for human-like voices
So-VITS-SVC
Python • MIT
Voice conversion — sing like any singer, speak like any voice
Bark (Suno)
Python • MIT
Text-to-speech — generate audio, music, and sound effects
Whisper (OpenAI)
Python • MIT
Speech recognition — transcribe 99 languages with near-human accuracy
F5-TTS
Python • Apache-2.0
Natural text-to-speech — flow matching for human-like voices
Retrieval-based Voice Conversion
Python • MIT
Real-time voice conversion — low latency, high quality
AudioCraft (Meta)
Python • MIT
Music generation — Meta open-source models for audio creation
Bark (Suno)
Python • MIT
Text-to-speech — generate audio, music, and sound effects
Whisper (OpenAI)
Python • MIT
Speech recognition — transcribe 99 languages with near-human accuracy