Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
☆10,303Mar 25, 2026Updated 6 months ago
Alternatives and similar repositories for Amphion
Users that are interested in Amphion are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SOTA Open Source TTS☆32,918Sep 16, 2026Updated 2 weeks ago
- Instant voice cloning by MIT and MyShell. Audio foundation model.☆37,750Apr 19, 2025Updated last year
- Inference and training library for high-quality TTS models.☆5,591Dec 10, 2024Updated last year
- Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.☆23,818May 25, 2026Updated 4 months ago
- Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"☆15,317Sep 21, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching☆1,362Updated this week
- zero-shot voice conversion & singing voice conversion, with real-time support☆3,892Apr 20, 2025Updated last year
- 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production☆46,097Aug 16, 2024Updated 2 years ago
- Zero-Shot Speech Editing and Text-to-Speech in the Wild☆8,576May 30, 2026Updated 4 months ago
- Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor…☆23,659Mar 3, 2026Updated 7 months ago
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models☆6,363Aug 10, 2024Updated 2 years ago
- PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html☆2,220Sep 10, 2025Updated last year
- Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis☆1,164Aug 29, 2026Updated last month
- 🔊 Text-Prompted Generative Audio Model☆39,273Aug 19, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.☆1,869Jul 16, 2026Updated 2 months ago
- Foundational model for human-like, expressive TTS☆4,205Jul 30, 2024Updated 2 years ago
- Text-to-Audio/Music Generation☆2,642Sep 29, 2024Updated 2 years ago
- Implementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in Pytorch☆1,333Sep 24, 2023Updated 3 years ago
- Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audi…☆11,168Sep 9, 2026Updated 3 weeks ago
- EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine☆8,536Sep 3, 2026Updated last month
- Official PyTorch implementation of BigVGAN (ICLR 2023)☆1,229Sep 5, 2024Updated 2 years ago
- The official implementation of HierSpeech++☆1,237Feb 20, 2024Updated 2 years ago
- 1 min voice data can also be used to train a good TTS model! (few shot voice cloning)☆62,241Aug 18, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Generative models for conditional audio generation☆3,872Sep 18, 2026Updated 2 weeks ago
- AcademiCodec: An Open Source Audio Codec Model for Academic Research☆675Dec 27, 2023Updated 2 years ago
- Foundational Models for State-of-the-Art Speech and Text Translation☆11,885Sep 8, 2026Updated 3 weeks ago
- GLM-4-Voice | 端到端中英语音对话模型☆3,238Dec 5, 2024Updated last year
- The Open Source Code of UniAudio☆608Jul 22, 2024Updated 2 years ago
- FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music gener…☆449Jan 25, 2024Updated 2 years ago
- An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Spe…☆4,535Aug 14, 2025Updated last year
- An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/☆7,922Feb 11, 2024Updated 2 years ago
- A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone☆26,485Sep 8, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.☆7,651Dec 24, 2024Updated last year
- A generative speech model for daily dialogue.☆39,886Apr 10, 2026Updated 5 months ago
- An Open-Sourced LLM-empowered Foundation TTS System☆926Sep 28, 2025Updated last year
- Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice☆535Dec 22, 2025Updated 9 months ago
- Next-generation TTS model using flow-matching and DiT, inspired by Stable Diffusion 3☆439Sep 13, 2024Updated 2 years ago
- This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples a…☆665Jun 9, 2024Updated 2 years ago
- AudioLDM: Generate speech, sound effects, music and beyond, with text.☆2,911Jun 25, 2025Updated last year