A demo-level low-latency, high-throughput inference engine for whisper
☆20Nov 9, 2025Updated 9 months ago
Alternatives and similar repositories for nano-whisper
Users that are interested in nano-whisper are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆16Aug 4, 2026Updated last week
- wav2vec2 audio classification for prosodic boundary detection and other tasks☆42Aug 11, 2023Updated 2 years ago
- A high-performance batch audio transcription tool using nvidia/parakeet-tdt-0.6b-v2 to generate accurate, well-segmented SRT subtitles, w…☆18Dec 9, 2025Updated 8 months ago
- Compute WER and SER for speech recognition evaluation☆28Jun 6, 2026Updated 2 months ago
- ☆10Sep 25, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Multispeaker Community Vocoder Model for DiffSinger☆39Aug 11, 2025Updated 11 months ago
- Finetune Nemo parakeet ASR model with new language (support 8 bit optimizer). Experimental birwkv-fastconformer TDT for long-form ASR(8.5…☆26Nov 27, 2025Updated 8 months ago
- Code for InterSpeech 2024 Paper: LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition☆19Jul 16, 2024Updated 2 years ago
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- Taiwanese Hokkien Transliterator and Tokeniser☆50Jul 1, 2026Updated last month
- Dataset for Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models in Interspeech 2024.☆16Jul 4, 2024Updated 2 years ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Graph model execution API for Candle☆18Jul 27, 2025Updated last year
- ☆16Jun 25, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The case study and multilingfual performance of ICASSP submission☆24Sep 24, 2022Updated 3 years ago
- ☆12Sep 18, 2024Updated last year
- 一款隐身于 Mac 摄像头下方的智能提词器:专为视频录制、直播与会议设计,帮您保持自然眼神交流。支持苹果自带语音识别与本地 AI 大模型,能随着您的真实语速自动跟踪和滚动文案,彻底告别忘词与手动滑屏的烦恼。☆19Feb 24, 2026Updated 5 months ago
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- Voice conversion with just linear regression.☆37Sep 25, 2025Updated 10 months ago
- Running LLaMA 3 with Rust.☆10May 21, 2024Updated 2 years ago
- Script to generate VAD dataset used in Asteroid recipe☆21Sep 30, 2021Updated 4 years ago
- A Rust crate offering similar functionality to the Python transformers package using Candle.☆15Nov 19, 2024Updated last year
- ☆10Oct 8, 2021Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 4G GPU & 10 Minutes for train☆12Aug 9, 2023Updated 3 years ago
- ☆12Jun 14, 2024Updated 2 years ago
- ☆24Aug 1, 2026Updated last week
- MeloTTS demo on Axera☆14Jul 1, 2026Updated last month
- A docker image for One Student One Chip's debug exam☆10Sep 22, 2023Updated 2 years ago
- Sampling techniques for Candle.☆21Apr 3, 2024Updated 2 years ago
- This is a project of Interspeech2021 paper "SpecMix : A Mixed Sample Data Augmentation method for Training with Time-Frequency Domain Fea…☆11Sep 27, 2022Updated 3 years ago
- Script to demonstrate how to use a Language Model for Semantic Turn Detection. Refer to blog post for full details.☆19May 9, 2025Updated last year
- The official repository for AdaMuon☆39Aug 27, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment☆16Apr 13, 2022Updated 4 years ago
- A chinese singing voice dataset, professional male singer, 105 songs, 132 minutes☆12Oct 19, 2023Updated 2 years ago
- The source code for the paper CrossSinger (asru2023)☆18Oct 12, 2023Updated 2 years ago
- Digital Speech Processing in PyTorch.☆15Aug 12, 2022Updated 3 years ago
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago
- Automatically Update LLM Papers Daily using Github Actions. Ref: https://github.com/Vincentqyw/cv-arxiv-daily☆10Aug 3, 2026Updated last week
- This is the accompanying repository to the paper - Automatic Estimation of Singing Voice Musical Dynamics☆16Oct 28, 2024Updated last year