A demo-level low-latency, high-throughput inference engine for whisper
☆20Nov 9, 2025Updated 8 months ago
Alternatives and similar repositories for nano-whisper
Users that are interested in nano-whisper are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆15Jun 9, 2026Updated last month
- wav2vec2 audio classification for prosodic boundary detection and other tasks☆42Aug 11, 2023Updated 2 years ago
- A high-performance batch audio transcription tool using nvidia/parakeet-tdt-0.6b-v2 to generate accurate, well-segmented SRT subtitles, w…☆18Dec 9, 2025Updated 7 months ago
- Compute WER and SER for speech recognition evaluation☆27Jun 6, 2026Updated last month
- ☆10Sep 25, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Multispeaker Community Vocoder Model for DiffSinger☆39Aug 11, 2025Updated 11 months ago
- Finetune Nemo parakeet ASR model with new language (support 8 bit optimizer). Experimental birwkv-fastconformer TDT for long-form ASR(8.5…☆26Nov 27, 2025Updated 7 months ago
- Code for InterSpeech 2024 Paper: LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition☆19Jul 16, 2024Updated 2 years ago
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- Taiwanese Hokkien Transliterator and Tokeniser☆49Jul 1, 2026Updated 2 weeks ago
- Dataset for Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models in Interspeech 2024.☆16Jul 4, 2024Updated 2 years ago
- Script to generate VAD dataset used in Asteroid recipe☆21Sep 30, 2021Updated 4 years ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Graph model execution API for Candle☆18Jul 27, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆16Jun 25, 2024Updated 2 years ago
- The case study and multilingfual performance of ICASSP submission☆24Sep 24, 2022Updated 3 years ago
- FastAPI Server Implementation for Bilibili Index TTS☆25Apr 13, 2025Updated last year
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- 一款隐身于 Mac 摄像头下方的智能提词器:专为视频录制、直播与会议设计,帮您保持自然眼神交流。支持苹果自带语音识别与本地 AI 大模型,能随着您的真实语速自动跟踪和滚动文案,彻底告别忘词与手动滑屏的烦恼。☆15Feb 24, 2026Updated 4 months ago
- ☆25Nov 25, 2025Updated 7 months ago
- Voice conversion with just linear regression.☆37Sep 25, 2025Updated 9 months ago
- Serving Next Generation Experimental Tracking for Machine Learning Operations☆15Mar 5, 2026Updated 4 months ago
- [ICMR 2025] Official Repository for The Paper, Let Network Decide What to Learn: Symbolic Music Understanding Model Based on Large-scale …☆18Aug 17, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Running LLaMA 3 with Rust.☆10May 21, 2024Updated 2 years ago
- A Rust crate offering similar functionality to the Python transformers package using Candle.☆15Nov 19, 2024Updated last year
- ☆12Jun 14, 2024Updated 2 years ago
- MeloTTS demo on Axera☆14Jul 1, 2026Updated 2 weeks ago
- ☆23Jul 12, 2026Updated last week
- Sampling techniques for Candle.☆21Apr 3, 2024Updated 2 years ago
- This is a project of Interspeech2021 paper "SpecMix : A Mixed Sample Data Augmentation method for Training with Time-Frequency Domain Fea…☆11Sep 27, 2022Updated 3 years ago
- 适用于 diffsinger 的多功能工具集☆11Apr 2, 2023Updated 3 years ago
- java implementation of Bert Tokenizer, support output onnx tensor for onnx model inference☆14Sep 4, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Script to demonstrate how to use a Language Model for Semantic Turn Detection. Refer to blog post for full details.☆18May 9, 2025Updated last year
- The official repository for AdaMuon☆39Aug 27, 2025Updated 10 months ago
- NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment☆16Apr 13, 2022Updated 4 years ago
- ☆13Sep 12, 2024Updated last year
- A chinese singing voice dataset, professional male singer, 105 songs, 132 minutes☆12Oct 19, 2023Updated 2 years ago
- The source code for the paper CrossSinger (asru2023)☆18Oct 12, 2023Updated 2 years ago
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago