Official implementation of "WhisperNER: Unified Open Named Entity and Speech Recognition"
☆199Feb 25, 2025Updated last year
Alternatives and similar repositories for whisper-ner
Users that are interested in whisper-ner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Multi-task model for named-entity recognition, relation extraction, entity mention detection and coreference resolution.☆46Jun 26, 2024Updated 2 years ago
- Whisper with Medusa heads☆858Jul 2, 2026Updated 2 months ago
- Drax: Speech Recognition with Discrete Flow Matching☆75Oct 15, 2025Updated 11 months ago
- A curated list of awesome papers on contextualizing E2E ASR outputs☆82May 10, 2023Updated 3 years ago
- Code for ICASSP 2024 Paper: RECAP: Retrieval-Augmented Audio Captioning☆16Jun 23, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 🍒 Dynamically inline assets into the DOM using Fetch Injection. Mirror of Fetch Inject on Codeberg.☆14May 26, 2024Updated 2 years ago
- ☆13Jun 1, 2026Updated 3 months ago
- ANDROID APP that can RECOGNIZE VLC LIVE AUDIO/VIDEO STREAMING (using free Android Developers Speech Recognition API) then TRANSLATE (usin…☆15Aug 27, 2026Updated 3 weeks ago
- Code for the paper: How Much Context Does My Attention-Based ASR System Need?☆13Jul 29, 2026Updated last month
- Versatile App Creation: Build ChatGPT apps for mobile or desktop☆19Feb 27, 2024Updated 2 years ago
- CTC decoder with hotwords for ASR.☆41Sep 9, 2026Updated 2 weeks ago
- PANiC - PAraphrasing Noun-Compounds☆15Apr 6, 2018Updated 8 years ago
- Bulk unsubscribe from emails in your gmail account☆25Jun 15, 2024Updated 2 years ago
- Acoustic Neighbor Embeddings☆33Sep 11, 2026Updated last week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot Classification☆11Aug 12, 2023Updated 3 years ago
- Fully neural approach for text chunking☆420Oct 23, 2025Updated 11 months ago
- [Interspeech 2024] Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation☆16Nov 28, 2024Updated last year
- Task-based Agentic Framework using StrictJSON as the core☆462Jun 8, 2026Updated 3 months ago
- Things you can do with the token embeddings of an LLM☆1,450Dec 1, 2025Updated 9 months ago
- first base model for full-duplex conversational audio☆1,800Jan 5, 2025Updated last year
- ☆37May 20, 2022Updated 4 years ago
- ☆89Jul 31, 2025Updated last year
- INTERSPEECH 23 - Refunction Whisper to recognize new tasks with adapters!☆41Sep 11, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.☆4,117Jan 8, 2025Updated last year
- Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks https://arxiv.org/abs/19…☆14Apr 16, 2020Updated 6 years ago
- Transcribing Speech with Multinomial Diffusion, training code and models.☆80Sep 27, 2023Updated 2 years ago
- An OpenAI-compatible ASR/STT API server powered by Meta's omnilingual-asr model. Supports real-time streaming via WebSocket and batch tra…☆20Jan 2, 2026Updated 8 months ago
- Code for our paper: *Shamsian, *Kleinfeld, Globerson & Chechik, "Learning Object Permanence from Video"☆68Nov 20, 2024Updated last year
- [EMNLP Main '25] LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation☆158May 18, 2025Updated last year
- The creative suite for character-driven AI experiences.☆192Sep 6, 2024Updated 2 years ago
- Custom firmware for Realtek RTL8761B* adapters. From research "Reverse engineering Realtek RTL8761B* Bluetooth chips, to make better Blue…☆22May 3, 2026Updated 4 months ago
- Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces☆11,136Aug 31, 2026Updated 3 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Faster Learned Sparse Retrieval with Block-Max Pruning. ACM SIGIR 2024.☆37Jan 14, 2026Updated 8 months ago
- An electron Wrapper for Open-Interpreter for the lablab.ai hackathon☆12Oct 14, 2023Updated 2 years ago
- An Open Source text-to-speech system built by inverting Whisper.☆4,650Dec 14, 2025Updated 9 months ago
- Use DEMUCS to split songs into multiple sources☆20Apr 11, 2022Updated 4 years ago
- Promting Whisper for Audio-Visual Speech Recognition, Code-Switched Speech Recognition, and Zero-Shot Speech Translation☆151Jan 16, 2024Updated 2 years ago
- ☆32Dec 2, 2020Updated 5 years ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 5 months ago