Official implementation of "WhisperNER: Unified Open Named Entity and Speech Recognition"
☆199Feb 25, 2025Updated last year
Alternatives and similar repositories for whisper-ner
Users that are interested in whisper-ner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Whisper with Medusa heads☆861Jul 2, 2026Updated last month
- Zero-shot Domain-sensitive Speech Recognition with Prompt-conditioning Fine-tuning (ASRU2023)☆26Oct 10, 2023Updated 2 years ago
- Drax: Speech Recognition with Discrete Flow Matching☆75Oct 15, 2025Updated 9 months ago
- A curated list of awesome papers on contextualizing E2E ASR outputs☆81May 10, 2023Updated 3 years ago
- A P2P blog and P2P Chat with no signalling server. Nothin' but RTC!☆16Nov 17, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆20Jun 3, 2024Updated 2 years ago
- Code for the paper: How Much Context Does My Attention-Based ASR System Need?☆12Jul 29, 2026Updated last week
- A fork of Lyra (version 1) that supports a webassembly build. See https://github.com/mayitayew/soundstream-wasm for a more recent version…☆25Jul 19, 2022Updated 4 years ago
- CTC decoder with hotwords for ASR.☆39Updated this week
- PANiC - PAraphrasing Noun-Compounds☆15Apr 6, 2018Updated 8 years ago
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot Classification☆11Aug 12, 2023Updated 2 years ago
- Fully neural approach for text chunking☆419Oct 23, 2025Updated 9 months ago
- Task-based Agentic Framework using StrictJSON as the core☆462Jun 8, 2026Updated 2 months ago
- Things you can do with the token embeddings of an LLM☆1,449Dec 1, 2025Updated 8 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICASSP2023] Source code, model links and open test sets for paper SeACo-Paraformer.☆44Mar 15, 2024Updated 2 years ago
- first base model for full-duplex conversational audio☆1,798Jan 5, 2025Updated last year
- ☆37May 20, 2022Updated 4 years ago
- INTERSPEECH 23 - Refunction Whisper to recognize new tasks with adapters!☆41Sep 11, 2023Updated 2 years ago
- Python tools for WhisperKit: Model conversion, optimization and evaluation☆246Mar 25, 2026Updated 4 months ago
- Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.☆4,101Jan 8, 2025Updated last year
- Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks https://arxiv.org/abs/19…☆14Apr 16, 2020Updated 6 years ago
- Transcribing Speech with Multinomial Diffusion, training code and models.☆80Sep 27, 2023Updated 2 years ago
- An OpenAI-compatible ASR/STT API server powered by Meta's omnilingual-asr model. Supports real-time streaming via WebSocket and batch tra…☆18Jan 2, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- How to take a picture with Hololens 2 and use the image as texture to create a photo you can hold in AR☆13Apr 21, 2020Updated 6 years ago
- The creative suite for character-driven AI experiences.☆192Sep 6, 2024Updated last year
- [ICASSP 2022] AISHELL-NER: Named Entity Recognition from Chinese Speech☆26Apr 20, 2022Updated 4 years ago
- Custom firmware for Realtek RTL8761B* adapters. From research "Reverse engineering Realtek RTL8761B* Bluetooth chips, to make better Blue…☆19May 3, 2026Updated 3 months ago
- An N-gram punctuator for Chinese and English.☆20Oct 14, 2025Updated 9 months ago
- mdast extension to parse and serialize MDX (or MDX.js)☆25Feb 10, 2024Updated 2 years ago
- Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces☆10,747Updated this week
- Faster Stable Diffusion using SSD-1B. A gradio app inside for demo.☆15Oct 26, 2023Updated 2 years ago
- Use DEMUCS to split songs into multiple sources☆20Apr 11, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆32Dec 2, 2020Updated 5 years ago
- Local realtime voice AI☆2,491Nov 26, 2025Updated 8 months ago
- ☆13Nov 11, 2023Updated 2 years ago
- Official PyTorch implementation for "MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens…☆48Jun 12, 2025Updated last year
- ☆58Feb 8, 2026Updated 6 months ago
- Provide Gradio custom components to make the diarization-based audio labeling process easier and faster.☆72Apr 22, 2026Updated 3 months ago
- [EMNLP 2025 Findings] A complete cross-modal RAG system for end-to-end speech-to-speech large models, including ASR-based Retrieval and E…☆31Jul 11, 2025Updated last year