Multi-speaker separation, identification, diarization ALL-IN-ONE. It can isolate the target speaker from a conversation audio and do ASR.
☆99Oct 13, 2025Updated 10 months ago
Alternatives and similar repositories for TargetDiarization
Users that are interested in TargetDiarization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Unofficial implementation of ConvNeXt-TTS powered by lightning☆18Oct 20, 2024Updated last year
- Attention-Based Encoder-Decoder Target-Speaker Voice Activity Detection for Robust Speaker Diarization☆32Sep 22, 2025Updated 11 months ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- A simple command line tool to calculate WER for ASR.☆14Jul 28, 2026Updated last month
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆35Sep 25, 2025Updated 11 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Adaptive Flow-Matching for Target Speaker Extraction☆43Jul 13, 2026Updated last month
- DUSTED: Spoken-Term Discovery using Discrete Speech Units☆17Oct 2, 2024Updated last year
- [WIP]Trying to implement "Ultra Low Complexity Deep Learning Based Noise Suppression." arXiv preprint arXiv:2312.08132 (2023).☆29May 29, 2024Updated 2 years ago
- Inference code for Interspeech 2025 paper, "LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec"☆37Oct 23, 2025Updated 10 months ago
- Grapheme-to-phoneme (G2P) conversion is the process of generating pronunciation for words based on their written form. It has a highly es…☆19Jun 14, 2021Updated 5 years ago
- Target Speaker Extraction Toolkit☆312Oct 4, 2025Updated 11 months ago
- ☆20Sep 2, 2024Updated 2 years ago
- System that ranks 2nd in DCASE 2022 Challenge Task 5: Few-shot Bioacoustic Event Detection☆28Jul 6, 2022Updated 4 years ago
- a kws demo on android☆40May 28, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Simple diarization model☆53Jun 13, 2025Updated last year
- A toolkit for speaker diarization.☆537Aug 4, 2026Updated last month
- Source code for the EMNLP 2025 paper “DM-Codec: Distilling Multimodal Representations for Speech Tokenization”☆57Jun 1, 2025Updated last year
- ☆12Nov 7, 2024Updated last year
- steps to perform text-based speaker diarization with kaldi toolkit☆12Nov 2, 2018Updated 7 years ago
- [EMNLP 2025 Findings] Official code for EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion☆43Sep 9, 2025Updated 11 months ago
- End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions☆94Nov 6, 2023Updated 2 years ago
- OpenSpeakerBeam-SS is an independent reimplementation of SpeakerBeam-SS, a real-time target speaker extraction model combining Conv-TasNe…☆19Aug 25, 2026Updated last week
- Implementation of Transfer Learning from Speaker Verification to Multi-speaker Text-To-Speech Synthesis (SV2TTS) in Persian language.☆12Oct 2, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An N-gram punctuator for Chinese and English.☆20Oct 14, 2025Updated 10 months ago
- Implementation of 'Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis', in MLX☆24Oct 30, 2024Updated last year
- Official Repository for "Efficient Vocal Source Separation Through Windowed RoFormer"☆46Oct 30, 2025Updated 10 months ago
- Python Wrapper of Silero VAD☆63May 8, 2025Updated last year
- Whisper Speech Quality Assessment (WhiSQA)☆16Apr 14, 2026Updated 4 months ago
- Understanding and Tackling Hallucinations in Large Audio-Language Models | ICASSP 2025, Interspeech 2024☆34Mar 14, 2025Updated last year
- Diffusion Singing Voice Conversion based on Grad-TTS from HuaWei☆173Oct 24, 2023Updated 2 years ago
- Persian Grapheme To Phoneme with Transformer in Pytorch☆11Sep 21, 2023Updated 2 years ago
- Kaldi style neural network training in pytorch for use in place of nnet3 in Kaldi.☆26Jul 25, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official repository of TF-Restormer for speech restoration☆17Updated this week
- silero-vad pytorch implement☆38Nov 23, 2024Updated last year
- ☆19Jan 6, 2025Updated last year
- Automatic speech annotator processing speech with voice activaty detection, overlapping speech detection, speaker diarization and automat…☆33Jun 14, 2024Updated 2 years ago
- A neural speech codec based on discrete WavLM representations☆26Aug 28, 2024Updated 2 years ago
- The Crepe plugin is an implementation of the CREPE monophonic pitch tracker, based on a deep convolutional neural network operating direc…☆16Dec 1, 2025Updated 9 months ago
- Silero VAD(ncnn): pre-trained enterprise-grade Voice Activity Detector.☆26Aug 21, 2024Updated 2 years ago