Ultra-Sortformer for Scalable Speaker Diarization
☆27Apr 9, 2026Updated 3 months ago
Alternatives and similar repositories for Ultra-Sortformer
Users that are interested in Ultra-Sortformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Meanflow and multilingual for F5-TTS model☆16Aug 23, 2025Updated 11 months ago
- StyleTTS 2 Optimized Training Fork☆32Feb 2, 2025Updated last year
- Pure-PyTorch Parakeet TDT inference☆51Mar 10, 2026Updated 4 months ago
- Native End-to-End Full-Duplex Spoken Language Model☆95Jul 30, 2026Updated last week
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval☆13Jun 27, 2025Updated last year
- A Weakly Supervised Forced Alignment for disluent speech☆15Nov 12, 2023Updated 2 years ago
- Code for Latent Speech-Text Transformer (LST)☆35Mar 12, 2026Updated 4 months ago
- ArtSpeech: Adaptive Text-to-Speech Synthesis with Articulatory Representations☆22Sep 21, 2025Updated 10 months ago
- Unofficial implementation of ConvNeXt-TTS powered by lightning☆18Oct 20, 2024Updated last year
- A collection of all our phonemeizers for dataset construction and inference☆30Feb 21, 2025Updated last year
- Forced alignment decoder for Whisper.☆16Mar 13, 2024Updated 2 years ago
- ☆45Apr 28, 2026Updated 3 months ago
- Official implementation of the paper "Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus" acc…☆77Jul 16, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The accompanying code for "Exploring the limits of decoder-only models trained on public speech recognition corpora" (Ankit Gupta, George…☆21Oct 11, 2024Updated last year
- ☆37Jun 9, 2026Updated last month
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆35Sep 25, 2025Updated 10 months ago
- ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models☆46Nov 18, 2025Updated 8 months ago
- AudioVisual Diarization - Supervised and Unsupervised☆15Nov 22, 2022Updated 3 years ago
- ☆33Jul 28, 2026Updated last week
- Train your own speech AI model from scratch☆152May 23, 2026Updated 2 months ago
- Multi-talker ASR based on DiCoW with Serialized Output Training☆21Sep 18, 2025Updated 10 months ago
- ☆37Jun 9, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- SpeechGLUE is a speech version of the GLUE benchmark, driven by text-to-speech.☆13Jun 2, 2023Updated 3 years ago
- LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models☆26Aug 11, 2024Updated last year
- An official implementation of Style-Talker for Spoken Dialogue Generation☆23Jan 12, 2025Updated last year
- ☆100Jan 28, 2026Updated 6 months ago
- This repository contains a series of works on diffusion-based speech tokenizers, including the official implementation of the paper: "TaD…☆198Jan 25, 2026Updated 6 months ago
- [ICLR 2026] StableToken: A state-of-the-art noise-robust semantic speech tokenizer featuring Voting-LFQ for resilient SpeechLLMs.☆33Feb 27, 2026Updated 5 months ago
- Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis☆27Mar 21, 2025Updated last year
- ☆64Dec 24, 2025Updated 7 months ago
- DUSTED: Spoken-Term Discovery using Discrete Speech Units☆17Oct 2, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Simple tool for speech dataset augmentation for modeling various prosodies.☆14Jan 14, 2021Updated 5 years ago
- [INTERSPEECH 2025 Oral]Official code for "Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment"☆67Jun 16, 2025Updated last year
- [KDD 2026] Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe☆43Aug 10, 2025Updated 11 months ago
- Train no-reference speech quality estimators with multiple datasets via learned, per-dataset alignments.☆18Aug 1, 2025Updated last year
- DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors☆39Feb 11, 2025Updated last year
- ☆21Jul 15, 2024Updated 2 years ago
- A Diffrentiable WFST-based End-to-End Automatic Speech Recognition toollkit with flexible topology support☆12Feb 15, 2026Updated 5 months ago