Official implementation of "WhisperNER: Unified Open Named Entity and Speech Recognition"
☆199Feb 25, 2025Updated last year
Alternatives and similar repositories for whisper-ner
Users that are interested in whisper-ner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Multi-task model for named-entity recognition, relation extraction, entity mention detection and coreference resolution.☆46Jun 26, 2024Updated 2 years ago
- Whisper with Medusa heads☆860Jul 2, 2026Updated 2 weeks ago
- Zero-shot Domain-sensitive Speech Recognition with Prompt-conditioning Fine-tuning (ASRU2023)☆26Oct 10, 2023Updated 2 years ago
- A curated list of awesome papers on contextualizing E2E ASR outputs☆81May 10, 2023Updated 3 years ago
- Code for ICASSP 2024 Paper: RECAP: Retrieval-Augmented Audio Captioning☆16Jun 23, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🍒 Dynamically inline assets into the DOM using Fetch Injection. Mirror of Fetch Inject on Codeberg.☆15May 26, 2024Updated 2 years ago
- A P2P blog and P2P Chat with no signalling server. Nothin' but RTC!☆16Nov 17, 2023Updated 2 years ago
- ☆20Jun 3, 2024Updated 2 years ago
- ☆11Jun 1, 2026Updated last month
- ANDROID APP that can RECOGNIZE VLC LIVE AUDIO/VIDEO STREAMING (using free Android Developers Speech Recognition API) then TRANSLATE (usin…☆14May 5, 2024Updated 2 years ago
- CTC decoder with hotwords for ASR.☆38Jun 15, 2026Updated last month
- PANiC - PAraphrasing Noun-Compounds☆15Apr 6, 2018Updated 8 years ago
- Fully neural approach for text chunking☆416Oct 23, 2025Updated 8 months ago
- Task-based Agentic Framework using StrictJSON as the core☆462Jun 8, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [Interspeech 2024] Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation☆14Nov 28, 2024Updated last year
- Things you can do with the token embeddings of an LLM☆1,450Dec 1, 2025Updated 7 months ago
- [ICASSP2023] Source code, model links and open test sets for paper SeACo-Paraformer.☆44Mar 15, 2024Updated 2 years ago
- first base model for full-duplex conversational audio☆1,794Jan 5, 2025Updated last year
- ☆88Jul 31, 2025Updated 11 months ago
- INTERSPEECH 23 - Refunction Whisper to recognize new tasks with adapters!☆41Sep 11, 2023Updated 2 years ago
- Python tools for WhisperKit: Model conversion, optimization and evaluation☆245Mar 25, 2026Updated 3 months ago
- Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.☆4,091Jan 8, 2025Updated last year
- An OpenAI-compatible ASR/STT API server powered by Meta's omnilingual-asr model. Supports real-time streaming via WebSocket and batch tra…☆18Jan 2, 2026Updated 6 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- How to take a picture with Hololens 2 and use the image as texture to create a photo you can hold in AR☆13Apr 21, 2020Updated 6 years ago
- [EMNLP Main '25] LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation☆154May 18, 2025Updated last year
- The creative suite for character-driven AI experiences.☆191Sep 6, 2024Updated last year
- 2023年iThome鐵人賽「AI & Data」組佳作【30天內成為NLP大師:掌握關鍵工具和技巧】完整程式碼,該文章會從零開始教你該如何微調大型語言模型☆18Nov 21, 2024Updated last year
- Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces☆10,205Updated this week
- An N-gram punctuator for Chinese and English.☆18Oct 14, 2025Updated 9 months ago
- Faster Stable Diffusion using SSD-1B. A gradio app inside for demo.☆15Oct 26, 2023Updated 2 years ago
- An Open Source text-to-speech system built by inverting Whisper.☆4,625Dec 14, 2025Updated 7 months ago
- Promting Whisper for Audio-Visual Speech Recognition, Code-Switched Speech Recognition, and Zero-Shot Speech Translation☆151Jan 16, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆32Dec 2, 2020Updated 5 years ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 2 months ago
- Local realtime voice AI☆2,490Nov 26, 2025Updated 7 months ago
- ☆13Nov 11, 2023Updated 2 years ago
- Official PyTorch implementation for "MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens…☆48Jun 12, 2025Updated last year
- Official implementation of Auxiliary Learning by Implicit Differentiation [ICLR 2021]☆86Jul 25, 2024Updated last year
- [WACV 2023] Audio-Visual Efficient Conformer (AVEC) for Robust Speech Recognition☆101Feb 21, 2023Updated 3 years ago