An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
☆4,073Sep 6, 2026Updated this week
Alternatives and similar repositories for MOSS-TTS
Users that are interested in MOSS-TTS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A 1.6B causal Transformer audio tokenizer with streaming, variable bitrates, and semantic alignment across speech, sound, and music☆255Jun 16, 2026Updated 2 months ago
- A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation☆4,291Updated this week
- VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning☆36,857Updated this week
- A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning☆1,394Updated this week
- High-Quality Voice Cloning TTS for 600+ Languages☆10,636Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆1,318Aug 17, 2026Updated 3 weeks ago
- An open-source model for understanding speech, environmental sounds, and music through captioning, question answering, and reasoning☆657Updated this week
- Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streamin…☆13,325Mar 17, 2026Updated 5 months ago
- An open-weight 11B model series for long-form and real-time video understanding☆608Updated this week
- Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.☆13,775Jul 24, 2026Updated last month
- Ming-omni-tts: Simple and Efficient Unified Generation of Speech, Music, and Sound with Precise Control☆265Feb 26, 2026Updated 6 months ago
- A foundation model that generates synchronized video and audio in a single model☆1,110Updated this week
- SOTA Open Source TTS☆32,603Updated this week
- Open-Source Frontier Voice AI☆53,885Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official inference code for SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis☆953May 29, 2026Updated 3 months ago
- ☆572Apr 3, 2026Updated 5 months ago
- GLM-TTS: Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning☆1,063Apr 10, 2026Updated 4 months ago
- A TTS that fits in your CPU (and pocket)☆9,396Updated this week
- The open-source AI voice studio. Clone, dictate, create.☆52,646Aug 9, 2026Updated last month
- ☆216Jun 2, 2026Updated 3 months ago
- Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.☆23,516May 25, 2026Updated 3 months ago
- Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching☆1,056Dec 2, 2025Updated 9 months ago
- "ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"☆12,295Jul 29, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling☆211Jun 6, 2026Updated 3 months ago
- An end-to-end speech-to-speech language model that generates spoken responses without text guidance☆139Feb 13, 2026Updated 6 months ago
- A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics…☆972Apr 9, 2026Updated 5 months ago
- SoTA open-source TTS☆26,318Jul 21, 2026Updated last month
- An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System☆23,838Aug 18, 2026Updated 3 weeks ago
- A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness☆1,879Updated this week
- [ACL 2026 Main] Training, inference, and testing of the SAC speech codec model.☆110Nov 1, 2025Updated 10 months ago
- super expressive prompting model based on ltx2.3☆488May 23, 2026Updated 3 months ago
- On-device TTS model by Neuphonic☆6,273Jul 30, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation☆455Nov 27, 2025Updated 9 months ago
- Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a …☆5,614Updated this week
- Zonos2 is a leading open-weight text-to-speech MoE.☆311Jul 6, 2026Updated 2 months ago
- MiMo-Audio: Audio Language Models are Few-Shot Learners☆1,080Jun 17, 2026Updated 2 months ago
- Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.☆9,370Aug 26, 2026Updated last week
- 🌋LavaSR: Fast Speech restoration and enhancement☆587Jun 19, 2026Updated 2 months ago
- Code for 'JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion'☆267May 11, 2026Updated 3 months ago