☆16Jun 25, 2025Updated last year
Alternatives and similar repositories for VocalStory
Users that are interested in VocalStory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Sep 16, 2024Updated last year
- ☆28Sep 14, 2024Updated last year
- A Chinese Expressive Long-dialogue Speech Dataset with Scripts☆21Nov 11, 2024Updated last year
- A simple VAD method☆11May 27, 2019Updated 7 years ago
- ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations☆183Mar 6, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- end-to-end text to audio scene generation model☆50Jun 16, 2026Updated last month
- Official PyTorch implementation of the Interspeech 2023 paper☆29Jul 5, 2023Updated 3 years ago
- Curated list for papers, codes and resources related to Text-to-Audio (TTA) Generation☆74May 27, 2026Updated last month
- Torch implementation of ViT based classifier for Audio classification☆12May 22, 2022Updated 4 years ago
- Official Code for ParrotTTS☆58Oct 13, 2024Updated last year
- Official repository for the paper "Audio ControlNet for Fine-Grained Audio Generation and Editing".☆75Feb 7, 2026Updated 5 months ago
- Here we will track the latest Audio AI Agent, including speech, music, sound effects, etc.☆16Dec 8, 2023Updated 2 years ago
- This is the official code for ACM CIKM 2025 Paper: ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive …☆59Dec 21, 2025Updated 7 months ago
- Simple but Useful Layers based on Tensorflow☆14Mar 29, 2020Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆51Jun 25, 2025Updated last year
- Language independent SSL-based Speaker Anonymization system☆20May 28, 2024Updated 2 years ago
- ☆23Oct 17, 2024Updated last year
- Code for ICLR 2024 Paper: CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models☆23Jul 10, 2024Updated 2 years ago
- A pitch detection model trained to be robust against noise and reverberation environments.☆27Jan 21, 2025Updated last year
- [ACL 2026 Main] Open-Ended Speaking Style Modeling via Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training☆104Apr 6, 2026Updated 3 months ago
- FlowMirror-HydraVox — A natively accelerated multi-head autoregressive TTS system derived from CosyVoice 3.0. It predicts multiple tokens…☆49Feb 17, 2026Updated 5 months ago
- [AAAI 2024] Code for CTX-vec2wav in UniCATS☆130Jun 11, 2024Updated 2 years ago
- [KDD 2026] Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe☆42Aug 10, 2025Updated 11 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- official implementation of paper ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification☆14Mar 14, 2025Updated last year
- Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'☆162Mar 26, 2026Updated 3 months ago
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness (https://arxiv.org/abs/2511.07931)☆77Dec 23, 2025Updated 6 months ago
- Multi-Agent Optimization in Python☆22Nov 29, 2019Updated 6 years ago
- ☆47Apr 27, 2026Updated 2 months ago
- Official code for "F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization"☆169Mar 3, 2026Updated 4 months ago
- Ultra-low bitrate speech codec (0.27-1 kbps) with cross-modal alignment and real-time capabilities☆355Aug 27, 2025Updated 10 months ago
- Official Implementation of GLAP - General Language Audio Pretraining☆74May 14, 2026Updated 2 months ago
- This repository contains a series of works on diffusion-based speech tokenizers, including the official implementation of the paper: "TaD…☆198Jan 25, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SpEx+(tied) source code☆96Jul 6, 2023Updated 3 years ago
- Pytorch Implementation of the paper "M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis"☆122Dec 18, 2025Updated 7 months ago
- Histogram Layer Time Delay Neural Networks For Passive Sonar Classification☆19Jan 21, 2026Updated 6 months ago
- List of Podcast Feeds using iTunes API and script to download 6,000,000~ hours of English speech.☆31Apr 13, 2023Updated 3 years ago
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model☆25May 21, 2026Updated 2 months ago
- Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis☆27Mar 21, 2025Updated last year
- Pushing the Limits of Zero-shot End-to-End Speech Translation☆25Dec 12, 2024Updated last year