ποΈ Automatically transcribe audio/video into high-quality, speaker-specific Text-To-Speech datasets
β142Aug 10, 2025Updated 11 months ago
Alternatives and similar repositories for TTSizer
Users that are interested in TTSizer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Try to replicate the architecture of MiniMaxTTS mentioned in it's technical reportβ47Sep 2, 2025Updated 10 months ago
- β101Jan 19, 2026Updated 6 months ago
- poorman's ar-dit ttsβ45Dec 31, 2025Updated 6 months ago
- A Neural Audio Codec (NAC) for Universal Audioβ46May 30, 2025Updated last year
- pytorch model for contexless-phoneme prediction from speech audioβ32Oct 30, 2025Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This repository implement a novel zero-shot TTS framework, named Flamed-TTS, focusing on the efficient generation and dynamic pacing in β¦β57Aug 9, 2025Updated 11 months ago
- [ACL 2025] OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matchingβ45Feb 9, 2025Updated last year
- PyTorch implementation of Miipher-2 [2025] which is a speech restoration model by Google DeepMindβ70Sep 22, 2025Updated 9 months ago
- High quality text-to-speech based on StyleTTS 2.β78Apr 6, 2026Updated 3 months ago
- β16Nov 11, 2024Updated last year
- Inference for the STFT-VAE continuous audio codec (24kHz, 3.125Hz latent)β43Jul 12, 2026Updated last week
- This repo is text to speech with learnable audio encoder without alignment with transcript referenceβ54Sep 20, 2025Updated 10 months ago
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into oneβ26Aug 5, 2024Updated last year
- Train the next generation of TTS systems.β169Sep 13, 2024Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is the official code for ACM CIKM 2025 Paper: ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive β¦β59Dec 21, 2025Updated 7 months ago
- This repository contains a series of works on diffusion-based speech tokenizers, including the official implementation of the paper: "TaDβ¦β198Jan 25, 2026Updated 5 months ago
- A pitch detection model trained to be robust against noise and reverberation environments.β27Jan 21, 2025Updated last year
- Extract phoneme-level timestamps from speeh audio.β154Jun 7, 2026Updated last month
- β41Jul 15, 2025Updated last year
- β110Updated this week
- X-E-Speech: Joint Training Framework of Non-Autoregressive Cross-lingual Emotional Text-to-Speech and Voice Conversionβ112Apr 1, 2024Updated 2 years ago
- Unofficial PyTorch implementation of "Autoregressive Speech Synthesis without Vector Quantization (MELLE)"β41Jun 28, 2025Updated last year
- Variable Bitrate Residual Vector Quantization for Audio Codingβ54May 1, 2025Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The open source code of ALMTokenizer2: Towards Low bit-rate and Semantic-rich Audio Tokenizer with Flow-based Scalar Diffusion Transformeβ¦β45Sep 5, 2025Updated 10 months ago
- Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'β162Mar 26, 2026Updated 3 months ago
- [INTERSPEECH 2026 Oral]Official code for "Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis"β120Jun 21, 2026Updated last month
- Ultra-low-bitrate Speech Codec for Speech Language Modeling Applicationsβ92Dec 20, 2024Updated last year
- Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesisβ27Mar 21, 2025Updated last year
- β302Jul 22, 2025Updated 11 months ago
- This is the code for paper: XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecsβ96Sep 19, 2025Updated 10 months ago
- β17Jun 2, 2025Updated last year
- Official code for "F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization"β169Mar 3, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformersβ126Mar 20, 2025Updated last year
- Unofficial Pytorch implementation of SNAC: Speaker-normalized affine coupling layer in flow-based architecture for zero-shot multi-speakeβ¦β57Aug 7, 2023Updated 2 years ago
- [ACL 2026 Main] Training, inference, and testing of the SAC speech codec model.β108Nov 1, 2025Updated 8 months ago
- VyvoTTS: LLM-Based Text-to-Speech Training Frameworkβ257Apr 8, 2026Updated 3 months ago
- FlashCosyVoice: A lightweight vLLM implementation built from scratch for CosyVoice.β250Feb 25, 2026Updated 4 months ago
- text to speechβ10Mar 19, 2024Updated 2 years ago
- β26Sep 22, 2022Updated 3 years ago