Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
☆319Dec 22, 2025Updated 8 months ago
Alternatives and similar repositories for Whisper-Finetune
Users that are interested in Whisper-Finetune are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training wit…☆1,223May 8, 2026Updated 3 months ago
- BELLE: Be Everyone's Large Language model Engine(开源中文对话大模型)☆8,275Oct 16, 2024Updated last year
- (WIP)long form speech generatoins☆30Apr 2, 2025Updated last year
- Torch Audio Forced Aligner for Mixed Chinese (Mandarin or Cantonese) and English.☆62Sep 5, 2025Updated 11 months ago
- Dolphin is a multilingual, multitask ASR model jointly trained by DataoceanAI and Tsinghua University.☆784Jun 11, 2026Updated 2 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 基于 faster-whisper 的伪实时语音转写服务☆241Apr 29, 2025Updated last year
- Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR be…☆1,974Feb 25, 2026Updated 6 months ago
- Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio…☆9,179Updated this week
- Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenA…☆20,081Updated this week
- ☆861Jun 7, 2024Updated 2 years ago
- We Speech Transcript based on LLM, in 300 lines of code.☆182Jun 20, 2025Updated last year
- Text-to-text alignment algorithm for speech recognition error analysis.☆33Jun 23, 2026Updated 2 months ago
- Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.☆20Mar 12, 2026Updated 5 months ago
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆35Sep 25, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Pseudo Streaming SenseVoice with Hotwords☆470Jun 15, 2026Updated 2 months ago
- Text-To-Speech for NotebookLM☆39Jul 20, 2025Updated last year
- 🔥 语音合成(TTS),语音克隆教程: https://dataxujing.github.io/TTS-paper/#/☆12Oct 29, 2024Updated last year
- A streaming audio reader, processor, and writer built on top of soundfile, and PyAV (bindings for FFmpeg)☆39Updated this week
- MooER: Moore-threads Open Omni model for speech-to-speech intERaction. MooER-omni includes a series of end-to-end speech interaction mode…☆220Jan 8, 2025Updated last year
- Term Project at GTCMT exploring phase based features for Singing Voice Detection with Neural Networks☆11Apr 20, 2018Updated 8 years ago
- Compute WER and SER for speech recognition evaluation☆28Jun 6, 2026Updated 2 months ago
- ☆110Oct 16, 2025Updated 10 months ago
- Offline Speaker Diarization with SenseVoice by Sherpa ONNX.☆15Dec 23, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Port of Funasr's Sense-voice model in C/C++☆572Dec 19, 2025Updated 8 months ago
- Colab notebook for fine-tuning Qwen2-Audio with trl's SFT and PPO trainers.☆24Nov 23, 2024Updated last year
- A native-PyTorch library for large scale M-LLM (text/audio) training with tp/cp/dp.☆234Jul 2, 2026Updated last month
- ☆562Jul 10, 2024Updated 2 years ago
- A tool for calculating WER (Word Error Rate) in python.☆14Sep 18, 2024Updated last year
- The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.☆1,946Jul 5, 2024Updated 2 years ago
- Awesome Neural Codec Models, Text-to-Speech Synthesizers & Speech Language Models☆247Jul 9, 2026Updated last month
- Run different pipelines of WhisperX - Transcription, Diarization, VAD, Alignment completely OFFLINE.☆48Mar 30, 2025Updated last year
- We Speech Toolkit, LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction☆211Jul 17, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- OSUM & OSUM-EChat, open speech understanding model and empathetic spoken chatbot based on it, open-sourced by ASLP@NPU.☆497Nov 23, 2025Updated 9 months ago
- Official repository for the WenetSpeech-Chuan dataset.☆234Jul 14, 2026Updated last month
- Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.☆4,112Jan 8, 2025Updated last year
- Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching☆1,046Dec 2, 2025Updated 8 months ago
- Faster Whisper transcription with CTranslate2☆25,146Nov 19, 2025Updated 9 months ago
- 🤗 R1-AQA Model: mispeech/r1-aqa☆327Mar 28, 2025Updated last year
- Causal streaming adaptation of OpenAI Whisper for real-time transcription on small audio chunks.☆76Mar 31, 2026Updated 5 months ago