A demo-level low-latency, high-throughput inference engine for whisper
☆20Nov 9, 2025Updated 9 months ago
Alternatives and similar repositories for nano-whisper
Users that are interested in nano-whisper are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆17Updated this week
- wav2vec2 audio classification for prosodic boundary detection and other tasks☆42Aug 11, 2023Updated 3 years ago
- A high-performance batch audio transcription tool using nvidia/parakeet-tdt-0.6b-v2 to generate accurate, well-segmented SRT subtitles, w…☆18Dec 9, 2025Updated 8 months ago
- Compute WER and SER for speech recognition evaluation☆28Jun 6, 2026Updated 2 months ago
- ☆10Sep 25, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Multispeaker Community Vocoder Model for DiffSinger☆39Aug 11, 2025Updated last year
- Code for InterSpeech 2024 Paper: LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition☆19Jul 16, 2024Updated 2 years ago
- Chinese-Mimi 是对 Moshi 模型的声码器进行了中文语料上的适配。☆38Mar 13, 2025Updated last year
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- Dataset for Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models in Interspeech 2024.☆16Jul 4, 2024Updated 2 years ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Graph model execution API for Candle☆18Jul 27, 2025Updated last year
- ☆16Jun 25, 2024Updated 2 years ago
- The case study and multilingfual performance of ICASSP submission☆24Sep 24, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆12Sep 18, 2024Updated last year
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- 一款隐身于 Mac 摄像头下方的智能提词器:专为视频录制、直播与会议设计,帮您保持自然眼神交流。支持苹果自带语音识别与本地 AI 大模型,能随着您的真实语速自动跟踪和滚动文案,彻底告别忘词与手动滑屏的烦恼。☆19Feb 24, 2026Updated 6 months ago
- Port of Funasr's Paraformer model in C/C++☆43Jun 19, 2024Updated 2 years ago
- Voice conversion with just linear regression.☆37Sep 25, 2025Updated 11 months ago
- Serving Next Generation Experimental Tracking for Machine Learning Operations☆15Mar 5, 2026Updated 5 months ago
- Running LLaMA 3 with Rust.☆10May 21, 2024Updated 2 years ago
- Script to generate VAD dataset used in Asteroid recipe☆21Sep 30, 2021Updated 4 years ago
- 4G GPU & 10 Minutes for train☆12Aug 9, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Jun 14, 2024Updated 2 years ago
- ☆24Aug 1, 2026Updated 3 weeks ago
- Sampling techniques for Candle.☆21Apr 3, 2024Updated 2 years ago
- This is a project of Interspeech2021 paper "SpecMix : A Mixed Sample Data Augmentation method for Training with Time-Frequency Domain Fea…☆11Sep 27, 2022Updated 3 years ago
- 适用于 diffsinger 的多功能工具集☆10Apr 2, 2023Updated 3 years ago
- Script to demonstrate how to use a Language Model for Semantic Turn Detection. Refer to blog post for full details.☆19May 9, 2025Updated last year
- java implementation of Bert Tokenizer, support output onnx tensor for onnx model inference☆14Sep 4, 2023Updated 2 years ago
- The official repository for AdaMuon☆39Aug 27, 2025Updated last year
- NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment☆16Apr 13, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆12Sep 12, 2024Updated last year
- A chinese singing voice dataset, professional male singer, 105 songs, 132 minutes☆13Oct 19, 2023Updated 2 years ago
- The source code for the paper CrossSinger (asru2023)☆18Oct 12, 2023Updated 2 years ago
- Digital Speech Processing in PyTorch.☆15Aug 12, 2022Updated 4 years ago
- Automatically Update LLM Papers Daily using Github Actions. Ref: https://github.com/Vincentqyw/cv-arxiv-daily☆10Aug 24, 2026Updated last week
- This is the accompanying repository to the paper - Automatic Estimation of Singing Voice Musical Dynamics☆16Oct 28, 2024Updated last year
- We Speech Toolkit, LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction☆211Jul 17, 2026Updated last month