LoRA-based phoneme/prosody control for LLM-based TTS with no G2P - Lightweight adapter for edit and control the target language's phoneme-level pronunciation and prosody while preserving other languages' performance.
☆29Jul 8, 2026Updated last month
Alternatives and similar repositories for UtterTune
Users that are interested in UtterTune are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Evaluation tool used in the BigVSAN paper☆14Mar 22, 2024Updated 2 years ago
- Please visit https://thuhcsi.github.io/SnakeGAN/☆38Apr 25, 2023Updated 3 years ago
- Cantonese Grapheme-to-Phoneme Converter based on GitYCC/g2pW☆15Dec 10, 2024Updated last year
- ☆19Mar 22, 2024Updated 2 years ago
- narabas: Japanese phoneme forced alignment tool☆15Mar 15, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Alignment examples for Interspeech 2024☆28Jul 5, 2024Updated 2 years ago
- ESLTTS dataset☆16Feb 6, 2025Updated last year
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated 2 years ago
- Voice conversion with just linear regression.☆37Sep 25, 2025Updated 10 months ago
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"☆11Apr 10, 2025Updated last year
- [Interspeech 2025] Official implementation of "Training-Free Voice Conversion with Factorized Optimal Transport"☆45Sep 24, 2025Updated 10 months ago
- Speaker-aware CTC (SACTC) for multi-talker overlapped speech recognition.☆22May 26, 2025Updated last year
- Ultra-low-bitrate Speech Codec for Speech Language Modeling Applications☆92Dec 20, 2024Updated last year
- poorman's ar-dit tts☆45Dec 31, 2025Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This repo is text to speech with learnable audio encoder without alignment with transcript reference☆56Sep 20, 2025Updated 10 months ago
- [AAAI 2024] Code for CTX-vec2wav in UniCATS☆130Jun 11, 2024Updated 2 years ago
- Code for Latent Speech-Text Transformer (LST)☆35Mar 12, 2026Updated 5 months ago
- [AAAI 2024] CTX-txt2vec, the acoustic model in UniCATS☆64Nov 18, 2024Updated last year
- Implementation of SoundStorm built upon SpeechTokenizer.☆116Nov 2, 2023Updated 2 years ago
- Unofficial implementation of wavenext vocoder☆59Aug 28, 2024Updated last year
- Unofficial implementation of ConvNeXt-TTS powered by lightning☆18Oct 20, 2024Updated last year
- Implementation of Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt (NAACL'24).☆119Jan 26, 2025Updated last year
- Collection of scripts from mHuBERT-147.☆35Nov 19, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official code for "F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization"☆169Mar 3, 2026Updated 5 months ago
- ☆16Aug 11, 2026Updated last week
- [IEEE OJSP'26, IEEE SLT'24] "Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization"☆46Jul 31, 2026Updated 2 weeks ago
- Llama-Mimi is a speech language model that uses a unified tokenizer (Mimi) and a single Transformer decoder (Llama) to jointly model sequ…☆31Sep 20, 2025Updated 10 months ago
- ☆88Nov 1, 2022Updated 3 years ago
- SC-CNN: Effective Speaker Conditioning Method for Zero-Shot Multi-Speaker Text-to-Speech Systems☆39Nov 1, 2023Updated 2 years ago
- Variable Bitrate Residual Vector Quantization for Audio Coding☆55May 1, 2025Updated last year
- ☆41May 15, 2023Updated 3 years ago
- Reimplementation of Miipher☆30Aug 16, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Descript Audio Codec - VAE Variant (.dac-vae): High-Fidelity Audio Compression with Variational Autoencoder☆39Aug 30, 2025Updated 11 months ago
- ☆35Sep 6, 2025Updated 11 months ago
- SpeechGLUE is a speech version of the GLUE benchmark, driven by text-to-speech.☆13Jun 2, 2023Updated 3 years ago
- PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions☆86Oct 11, 2024Updated last year
- Code for vec2wav 2.0, a speech token vocoder for VC. Paper: https://arxiv.org/abs/2409.01995☆79Dec 3, 2024Updated last year
- AAAI 2025: Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model☆309Oct 12, 2025Updated 10 months ago
- ☆10Jun 11, 2024Updated 2 years ago