A TTS model capable of generating ultra-realistic dialogue in one pass.
☆222Feb 19, 2026Updated 5 months ago
Alternatives and similar repositories for dia-multilingual
Users that are interested in dia-multilingual are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆131Jul 25, 2025Updated last year
- ☆302Jul 22, 2025Updated last year
- ☆35Sep 6, 2025Updated 10 months ago
- SoTA open-source TTS☆136Jun 7, 2025Updated last year
- VoXtream is a Full-Stream Zero-shot TTS model with Extremely Low Latency and Speaking rate Control☆245May 30, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆16Jun 28, 2025Updated last year
- This project is to train an RWKV LLM for TTS generation which compatible to other TTS engine(like fish/cosy/chattts).☆101Oct 8, 2025Updated 9 months ago
- [INTERSPEECH 2025 Oral]Official code for "Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment"☆67Jun 16, 2025Updated last year
- [EMNLP 2025 Findings] Official code for EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion☆43Sep 9, 2025Updated 10 months ago
- The open source code for SimpleSpeech series☆147Oct 8, 2024Updated last year
- ☆40Apr 3, 2025Updated last year
- Streaming and Fine-tuning for Chatterbox TTS☆292Jun 15, 2025Updated last year
- Inference for the STFT-VAE continuous audio codec (24kHz, 3.125Hz latent)☆43Jul 12, 2026Updated last week
- VyvoTTS: LLM-Based Text-to-Speech Training Framework☆257Apr 8, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), suppor…☆352Mar 28, 2026Updated 3 months ago
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆19,358Nov 19, 2025Updated 8 months ago
- Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis☆360Jun 25, 2026Updated last month
- Unofficial fairseq-free PyTorch implementation of UTMOS (v1, 2022), matching the original system.☆35Jun 6, 2026Updated last month
- Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation☆67Jun 16, 2026Updated last month
- LEMAS‑TTS is a multilingual zero‑shot text‑to‑speech system, supporting 10 languages: Chinese English Spanish Russian French German Ital…☆101Mar 31, 2026Updated 3 months ago
- The demo page for ALMTokenizer☆59Apr 14, 2025Updated last year
- Unofficial implementation of NVIDIA P-Flow TTS paper☆228Dec 24, 2024Updated last year
- VLLM Port of the Chatterbox TTS model☆379Oct 18, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Implementation of TTS model based on NVIDIA P-Flow TTS Paper☆77Jul 13, 2026Updated last week
- Official implementation of the paper "Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus" acc…☆77Jul 16, 2023Updated 3 years ago
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers☆126Mar 20, 2025Updated last year
- [Interspeech 2025] Official implementation of "Training-Free Voice Conversion with Factorized Optimal Transport"☆45Sep 24, 2025Updated 10 months ago
- Ultra-low-bitrate Speech Codec for Speech Language Modeling Applications☆92Dec 20, 2024Updated last year
- [ICASSP 2025] "FLowHigh: Towards efficient and high-quality audio super-resolution with single-step flow matching"☆118Jan 17, 2025Updated last year
- A Weakly Supervised Forced Alignment for disluent speech☆15Nov 12, 2023Updated 2 years ago
- Official implementation of paper: Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis☆55Sep 20, 2025Updated 10 months ago
- 🎙️ Automatically transcribe audio/video into high-quality, speaker-specific Text-To-Speech datasets ✨☆18May 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is the official train-dev-test release of the Interspeech2024 Discrete Speech Representation Challenge.☆32Jan 26, 2024Updated 2 years ago
- faster inference☆27Jan 20, 2025Updated last year
- This repository contains a series of works on diffusion-based speech tokenizers, including the official implementation of the paper: "TaD…☆198Jan 25, 2026Updated 6 months ago
- KVAE-Audio: a continuous full-band audio waveform autoencoder☆101Updated this week
- Towards Human-Sounding Speech☆6,260Dec 5, 2025Updated 7 months ago
- Text-To-Speech for NotebookLM☆39Jul 20, 2025Updated last year
- kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization☆16Nov 7, 2025Updated 8 months ago