WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
☆23,468Jul 13, 2026Updated 3 weeks ago
Alternatives and similar repositories for whisperX
Users that are interested in whisperX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Faster Whisper transcription with CTranslate2☆24,795Nov 19, 2025Updated 8 months ago
- Robust Speech Recognition via Large-Scale Weak Supervision☆106,840Jul 28, 2026Updated last week
- Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker…☆10,386Updated this week
- Port of OpenAI's Whisper model in C/C++☆52,660Updated this week
- Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper☆5,616Feb 23, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆13,041Oct 25, 2025Updated 9 months ago
- Silero VAD: pre-trained enterprise-grade Voice Activity Detector☆9,889Jul 16, 2026Updated 3 weeks ago
- Multilingual Automatic Speech Recognition with word-level timestamps and confidence☆2,836Sep 9, 2025Updated 10 months ago
- 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production☆45,865Aug 16, 2024Updated last year
- Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.☆4,101Jan 8, 2025Updated last year
- SOTA Open Source TTS☆32,081Updated this week
- 🔊 Text-Prompted Generative Audio Model☆39,224Aug 19, 2024Updated last year
- Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenA…☆19,709Updated this week
- JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.☆4,683Apr 3, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Whisper realtime streaming for long speech-to-text transcription and translation☆3,661Nov 12, 2025Updated 8 months ago
- Transcription, forced alignment, and audio indexing with OpenAI's Whisper☆2,280May 30, 2026Updated 2 months ago
- Instant voice cloning by MIT and MyShell. Audio foundation model.☆37,104Apr 19, 2025Updated last year
- A PyTorch-based Speech Toolkit☆11,743Jun 15, 2026Updated last month
- A nearly-live implementation of OpenAI's Whisper.☆4,216Updated this week
- Foundational Models for State-of-the-Art Speech and Text Translation☆11,837Jul 28, 2026Updated last week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆88,441Updated this week
- LLM inference in C/C++☆122,994Updated this week
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models☆6,328Aug 10, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.☆69,683Updated this week
- A multi-voice TTS system trained with an emphasis on quality☆14,865Nov 19, 2024Updated last year
- An Open Source text-to-speech system built by inverting Whisper.☆4,629Dec 14, 2025Updated 7 months ago
- Fast inference engine for Transformer models☆4,614Updated this week
- The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails…☆55,818Updated this week
- User-friendly AI Interface (Supports Ollama, OpenAI API, ...)☆148,143Updated this week
- LlamaIndex is the leading document agent and OCR platform☆51,446Updated this week
- SoTA open-source TTS☆25,876Jul 21, 2026Updated 2 weeks ago
- ☆3,575Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor…☆23,541Mar 3, 2026Updated 5 months ago
- Convert PDF to markdown + JSON quickly with high accuracy☆38,516Updated this week
- The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.☆124,612Updated this week
- A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Auto…☆18,007Updated this week
- OCR, layout analysis, reading order, table recognition in 90+ languages☆21,222Jul 23, 2026Updated 2 weeks ago
- Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.☆20,828Updated this week
- A python package to build AI-powered real-time audio applications☆2,011Jun 19, 2026Updated last month