open-source audio datasets
☆158Sep 7, 2023Updated 2 years ago
Alternatives and similar repositories for audio-datasets
Users that are interested in audio-datasets are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pretrained spoken language classifiers from audio.☆10Jan 21, 2021Updated 5 years ago
- LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation☆80Feb 24, 2021Updated 5 years ago
- PyTorch Implementation of Google Brain's WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis☆68Aug 3, 2021Updated 4 years ago
- Dynamic Mixing For Speech Processing (mix-on-the-fly)☆22Jul 19, 2022Updated 4 years ago
- This is a list of speech tasks and datasets, which can provide training data for Generative AI, AIGC, AI model training, intelligent spee…☆83Jun 7, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- VAE modified from Descript Audio Codec, which replaces the RVQ with VAE☆92Apr 2, 2024Updated 2 years ago
- UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation☆76Aug 30, 2021Updated 4 years ago
- A collection of datasets for the purpose of emotion recognition/detection in speech.☆420Sep 30, 2024Updated last year
- Official implementation of the source-filter HiFiGAN vocoder☆275Jul 29, 2023Updated 3 years ago
- Big Impulse Response Dataset☆159Oct 19, 2022Updated 3 years ago
- Code of the paper "Low-Latency Speech Separation Guided Diarization for Telephone Conversations"☆15Dec 22, 2022Updated 3 years ago
- PyTorch Implementation of DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs☆349Feb 21, 2022Updated 4 years ago
- BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis☆238Jul 13, 2022Updated 4 years ago
- Voice activity detection and speaker gender segmentation audiovisual corpus☆16Jan 20, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A toolset for easy formant extraction and visualization from wav files and TTS models☆33Sep 2, 2022Updated 3 years ago
- An official repo for the paper "Adapting Language-Audio Models as Few-Shot Audio Learners"☆31May 31, 2023Updated 3 years ago
- Fast audio data augmentation in PyTorch. Inspired by audiomentations. Useful for deep learning.☆1,162Nov 24, 2025Updated 8 months ago
- 🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).☆2,212Jun 6, 2024Updated 2 years ago
- 5Hz Deep-Compression Speech VAE for AR-Diffusion and CALMs☆57Nov 19, 2025Updated 8 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- ☆41May 15, 2023Updated 3 years ago
- Whisper Speech Quality Assessment (WhiSQA)☆16Apr 14, 2026Updated 3 months ago
- ☆15Jan 9, 2026Updated 6 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- "Learning Discrete and Continuous Factors of Data via Alternating Disentanglement" accepted at ICML2019☆22Aug 22, 2019Updated 6 years ago
- Code for Novel View Acoustic Synthesis paper☆54Aug 14, 2023Updated 2 years ago
- Official implementation of "Avocodo: Generative Adversarial Network for Artifact-Free Vocoder" (AAAI2023)☆154Feb 1, 2023Updated 3 years ago
- ☆17Oct 26, 2018Updated 7 years ago
- [ICASSP 2024] StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations☆141Apr 27, 2024Updated 2 years ago
- ☆14Mar 25, 2023Updated 3 years ago
- Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual Knowledge☆21Jul 25, 2022Updated 4 years ago
- Speech enhancement by time-varying pitch-dependent filtering of harmonics☆27Jul 3, 2014Updated 12 years ago
- Unified Speech Language Model for paper "SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models"(ICLR 2024)☆152Sep 14, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆15May 9, 2022Updated 4 years ago
- download the vggsound dataset☆22Feb 22, 2022Updated 4 years ago
- Official implementation for Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning☆100Nov 20, 2024Updated last year
- Score Normalization for NIST 2019 Speaker Recognition Evaluation☆10Nov 8, 2019Updated 6 years ago
- ☆16Feb 19, 2026Updated 5 months ago
- Dataset release for Emotional TTS in Indian Accent☆41Mar 25, 2026Updated 4 months ago
- Official repository of DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech, ICASSP 2023☆260Jun 5, 2025Updated last year