MOSS-Audio-Tokenizer is a Causal Transformer-based audio tokenizer built on the CAT architecture. Trained on 3M hours of diverse audio, it supports streaming and variable bitrates, delivering SOTA reconstruction and strong performance in generation and understanding—serving as a unified interface for next-generation native audio language models.
☆248Jun 16, 2026Updated last month
Alternatives and similar repositories for MOSS-Audio-Tokenizer
Users that are interested in MOSS-Audio-Tokenizer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026 Main] Training, inference, and testing of the SAC speech codec model.☆108Nov 1, 2025Updated 8 months ago
- This is the code for paper: XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs☆97Sep 19, 2025Updated 10 months ago
- A unified tokenizer that is capable of both extracting semantic information and enabling high-fidelity audio reconstruction.☆145Sep 19, 2025Updated 10 months ago
- [ACL 2026 Main] Open-Ended Speaking Style Modeling via Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training☆104Apr 6, 2026Updated 3 months ago
- Official code for "WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling"☆62Jun 27, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆88Feb 24, 2026Updated 5 months ago
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆125Jun 4, 2025Updated last year
- A curated list of models, benchmarks, tools and guides for audio editing☆34Jul 7, 2026Updated 2 weeks ago
- Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-Step High-Fidelity Audio Generation☆145Mar 8, 2026Updated 4 months ago
- ☆24Nov 16, 2025Updated 8 months ago
- State-of-the-art continious audio tokenization☆40Mar 9, 2026Updated 4 months ago
- [ICASSP 2026] Official code for "Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration"☆17Apr 16, 2026Updated 3 months ago
- ☆35Sep 6, 2025Updated 10 months ago
- MOSS-Audio is an open-source foundation model for unified audio understanding, enabling speech, sound, music, captioning, QA, and reasoni…☆617Jun 2, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [INTERSPEECH 2026 Oral]Official code for "Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis"☆121Jun 21, 2026Updated last month
- ☆187Aug 25, 2025Updated 11 months ago
- This repository contains a series of works on diffusion-based speech tokenizers, including the official implementation of the paper: "TaD…☆198Jan 25, 2026Updated 6 months ago
- MOSS-Speech is a true speech-to-speech large language model without text guidance.☆138Feb 13, 2026Updated 5 months ago
- [ACL 2026 Main] MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows☆142Sep 2, 2025Updated 10 months ago
- Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice☆521Dec 22, 2025Updated 7 months ago
- Ming-omni-tts: Simple and Efficient Unified Generation of Speech, Music, and Sound with Precise Control☆263Feb 26, 2026Updated 5 months ago
- ☆44Apr 26, 2026Updated 3 months ago
- LongCat Audio Tokenizer and Detokenizer☆301May 9, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICASSP 2026] Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis☆40Dec 24, 2025Updated 7 months ago
- HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding☆39Jun 8, 2026Updated last month
- end-to-end text to audio scene generation model☆50Jun 16, 2026Updated last month
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation☆451Nov 27, 2025Updated 7 months ago
- Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"☆218Sep 19, 2024Updated last year
- This repository contains a series of works on diffusion-based speech tokenizers, including the official implementation of the paper: "TaD…☆77Jan 25, 2026Updated 6 months ago
- KVAE-Audio: a continuous full-band audio waveform autoencoder☆101Updated this week
- Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.☆275Jul 17, 2026Updated last week
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.☆142Apr 7, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆50Apr 5, 2026Updated 3 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- [ICML 2026] Prism: Spectral-Aware Block-Sparse Attention☆27May 22, 2026Updated 2 months ago
- 5Hz Deep-Compression Speech VAE for AR-Diffusion and CALMs☆57Nov 19, 2025Updated 8 months ago
- WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling☆210Jun 6, 2026Updated last month
- MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flex…☆1,362Mar 23, 2026Updated 4 months ago
- We Speech Toolkit, LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction☆206Jul 17, 2026Updated last week