GPT-4o-level, real-time spoken dialogue system.
☆375Jan 27, 2025Updated last year
Alternatives and similar repositories for SpeechGPT-2.0-preview
Users that are interested in SpeechGPT-2.0-preview are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis☆363Jun 25, 2026Updated 2 months ago
- Ultra-low-bitrate Speech Codec for Speech Language Modeling Applications☆92Dec 20, 2024Updated last year
- MOSS-Speech is a true speech-to-speech large language model without text guidance.☆139Feb 13, 2026Updated 6 months ago
- This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples a…☆663Jun 9, 2024Updated 2 years ago
- ✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM☆395May 27, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 5Hz Deep-Compression Speech VAE for AR-Diffusion and CALMs☆57Nov 19, 2025Updated 9 months ago
- The open source code for SimpleSpeech series☆147Oct 8, 2024Updated last year
- Text-To-Speech for NotebookLM☆39Jul 20, 2025Updated last year
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction☆222Feb 28, 2025Updated last year
- This is the code for paper: XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs☆97Sep 19, 2025Updated 11 months ago
- Unified Speech Language Model for paper "SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models"(ICLR 2024)☆152Sep 14, 2023Updated 2 years ago
- Real-time Speech-Text Foundation Model Toolkit (wip)☆255Mar 26, 2025Updated last year
- A Survey of Spoken Dialogue Models (60 pages)☆316Nov 28, 2024Updated last year
- MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flex…☆1,391Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆193Aug 25, 2025Updated last year
- Make-An-Audio-3: Transforming Text/Video into Audio via Flow-based Large Diffusion Transformers☆121May 19, 2025Updated last year
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆127Jun 4, 2025Updated last year
- MOSS-Audio-Tokenizer is a Causal Transformer-based audio tokenizer built on the CAT architecture. Trained on 3M hours of diverse audio, i…☆254Jun 16, 2026Updated 2 months ago
- Please visit https://thuhcsi.github.io/SnakeGAN/