Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
☆1,220May 8, 2026Updated 4 months ago
Alternatives and similar repositories for Whisper-Finetune
Users that are interested in Whisper-Finetune are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training wit…☆319Dec 22, 2025Updated 8 months ago
- ☆562Jul 10, 2024Updated 2 years ago
- [WIP] Scripts for fine-tuning Whisper☆221Jul 2, 2026Updated 2 months ago
- Production First and Production Ready End-to-End Speech Recognition Toolkit☆5,239Sep 7, 2026Updated last week
- Fine-tune and evaluate Whisper models for Automatic Speech Recognition (ASR) on custom datasets or datasets from huggingface.☆368May 23, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The repo provides information about KeSpeech Mandarin dialect dataset.☆186Oct 13, 2022Updated 3 years ago
- We Speech Transcript based on LLM, in 300 lines of code.☆182Jun 20, 2025Updated last year
- chinese speech pretrained models☆1,214Aug 23, 2024Updated 2 years ago
- Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR be…☆1,993Feb 25, 2026Updated 6 months ago
- Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenA…☆20,443Updated this week
- ☆1,502Jul 16, 2026Updated 2 months ago
- The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.☆1,953Jul 5, 2024Updated 2 years ago
- This is a repository for fine-tuning Qwen2-Audio, currently supporting Distributed Data Parallel (DDP) and DeepSpeed.☆50Jul 28, 2025Updated last year
- A Framework for Speech, Language, Audio, Music Processing with Large Language Model☆1,064Jan 15, 2026Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio…☆9,334Sep 10, 2026Updated last week
- SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.☆557Mar 29, 2025Updated last year
- Faster Whisper transcription with CTranslate2☆25,478Nov 19, 2025Updated 10 months ago
- ☆89Jul 31, 2025Updated last year
- Text Normalization & Inverse Text Normalization☆824Jul 29, 2026Updated last month
- SALMONN family: A suite of advanced multi-modal LLMs☆1,535Aug 24, 2026Updated 3 weeks ago
- Chinese text normalization for speech processing☆740Mar 18, 2023Updated 3 years ago
- Update ASR paper everyday☆514May 16, 2026Updated 4 months ago
- Dolphin is a multilingual, multitask ASR model jointly trained by DataoceanAI and Tsinghua University.☆790Jun 11, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆860Jun 7, 2024Updated 2 years ago
- A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization☆3,144Dec 8, 2025Updated 9 months ago
- The dataset of Speech Recognition☆469Jan 4, 2026Updated 8 months ago
- Production First and Production Ready End-to-End Text-to-Speech Toolkit☆417Nov 20, 2025Updated 10 months ago
- Silero VAD: pre-trained enterprise-grade Voice Activity Detector☆10,262Updated this week
- Whisper realtime streaming for long speech-to-text transcription and translation☆3,671Nov 12, 2025Updated 10 months ago
- 中文标点符号模型,可以给文本添加标点符号。☆145Dec 24, 2024Updated last year
- A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/…☆687Jun 2, 2026Updated 3 months ago
- Pseudo Streaming SenseVoice with Hotwords☆470Jun 15, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- FSA/FST algorithms, differentiable, with PyTorch compatibility.☆1,360Jul 11, 2026Updated 2 months ago
- An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Spe…☆4,510Aug 14, 2025Updated last year
- End-to-End Speech Processing Toolkit☆9,963Updated this week
- A curated list of awesome papers on contextualizing E2E ASR outputs☆82May 10, 2023Updated 3 years ago
- Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.☆4,115Jan 8, 2025Updated last year
- ✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM☆397May 27, 2025Updated last year
- A ctc decoder for both online and offline asr model☆66Nov 18, 2023Updated 2 years ago