☆37Jun 9, 2025Updated last year
Alternatives and similar repositories for SparkVox
Users that are interested in SparkVox are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A large-scale speech corpus introduced in Spark-TTS, built from diverse open-source datasets for training text-to-speech (TTS) systems.☆115May 5, 2025Updated last year
- Scaled diffusion transformer for text-to-speech synthesis (DiT + T5Gemma2 conditioning, TorchTitan & Megatron backends, tested up to 1024…☆24Mar 29, 2026Updated 3 months ago
- ☆35Sep 6, 2025Updated 10 months ago
- The official repository for the paper “NonVerbalSpeech-38K: A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understandi…☆68Dec 26, 2025Updated 6 months ago
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"☆11Apr 10, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"☆218Sep 19, 2024Updated last year
- ☆22Jan 27, 2026Updated 5 months ago
- An AR+AR TTS attempt.☆18Jan 13, 2025Updated last year
- The open source code of ALMTokenizer2: Towards Low bit-rate and Semantic-rich Audio Tokenizer with Flow-based Scalar Diffusion Transforme…☆45Sep 5, 2025Updated 10 months ago
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆34Sep 25, 2025Updated 9 months ago
- Official Repository of Paper: "SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding" (IC…☆72Apr 27, 2026Updated 2 months ago
- Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-Step High-Fidelity Audio Generation☆144Mar 8, 2026Updated 4 months ago
- ☆101Jan 19, 2026Updated 6 months ago
- Descript Audio Codec - VAE Variant (.dac-vae): High-Fidelity Audio Compression with Variational Autoencoder☆38Aug 30, 2025Updated 10 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A pitch detection model trained to be robust against noise and reverberation environments.☆27Jan 21, 2025Updated last year
- Meanflow and multilingual for F5-TTS model☆16Aug 23, 2025Updated 10 months ago
- Official repo for CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations☆67Jan 16, 2025Updated last year
- [ACL 2026 Main] Training, inference, and testing of the SAC speech codec model.☆108Nov 1, 2025Updated 8 months ago
- finetune llm part for spark-tts model☆125Mar 25, 2025Updated last year
- Torch Audio Forced Aligner for Mixed Chinese (Mandarin or Cantonese) and English.☆61Sep 5, 2025Updated 10 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- Official Repository of Paper: "Emilia-NV: A Non-Verbal Speech Dataset with Word-Level Annotation for Human-Like Speech Modeling"☆91Sep 18, 2025Updated 10 months ago
- ☆44Jul 5, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An official implementation of Style-Talker for Spoken Dialogue Generation☆23Jan 12, 2025Updated last year
- ☆33Oct 28, 2025Updated 8 months ago
- 5Hz Deep-Compression Speech VAE for AR-Diffusion and CALMs☆57Nov 19, 2025Updated 8 months ago
- [ACL 2024] Generative Pre-Trained Speech Language Model with Efficient Hierarchical Transformer☆70Nov 1, 2024Updated last year
- Voxtral Codec : Combining Semantic VQ and Acoustic FSQ for Ultra-Low Bitrate Speech Generation (Voxtral TTS Backbone)☆15Mar 27, 2026Updated 3 months ago
- Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'☆162Mar 26, 2026Updated 3 months ago
- A TTS Trained on Universal Audio.☆41Jun 6, 2025Updated last year
- Event Relation in Text-to-Audio (TTA) Generation☆21Feb 26, 2025Updated last year
- [NeurIPS' 25] Benchmark for evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges.☆225Dec 9, 2025Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for Latent Speech-Text Transformer (LST)☆35Mar 12, 2026Updated 4 months ago
- Source code for the EMNLP 2025 paper “DM-Codec: Distilling Multimodal Representations for Speech Tokenization”☆57Jun 1, 2025Updated last year
- [INTERSPEECH 2024] The official implementation of EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for …☆181Updated this week
- Native full-duplex speech dialogue inference for BayLing-Duplex.☆63Jun 22, 2026Updated 3 weeks ago
- ☆35Oct 23, 2025Updated 8 months ago
- [INTERSPEECH 2025 Oral]Official code for "Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment"☆67Jun 16, 2025Updated last year
- LEMAS‑TTS is a multilingual zero‑shot text‑to‑speech system, supporting 10 languages: Chinese English Spanish Russian French German Ital…☆100Mar 31, 2026Updated 3 months ago