This repository is built with a focus on practical ways to obtain and work with the audio data of audioset. You can use this repository to download and precprocess audioset wav files for running the recipies of Audio Spectogram Transformer (AST) and Masked Autoencoder that listen (Audio - MAE).
☆17Jun 12, 2025Updated last year
Alternatives and similar repositories for Audioset
Users that are interested in Audioset are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Aug 23, 2024Updated last year
- ☆47Apr 2, 2025Updated last year
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆35Feb 11, 2026Updated 5 months ago
- ☆12Jun 17, 2017Updated 9 years ago
- A PyTorch Dataset for Slakh2100☆10Feb 14, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICML'24] Creative Text-to-Audio Generation via Synthesizer Programming☆41Sep 26, 2024Updated last year
- SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.☆123Jan 28, 2026Updated 6 months ago
- ☆53Mar 24, 2026Updated 4 months ago
- SongDriver2 achieves a balance between real-time emotion fit and soft transitions, enhancing the coherence of the generated music.☆29Nov 15, 2025Updated 8 months ago
- The repository of the paper: Wang et al., Learning interpretable representation for controllable polyphonic music generation, ISMIR 2020.☆45Mar 22, 2024Updated 2 years ago
- ☆40May 12, 2025Updated last year
- [CVPR 2025] Pytorch implementation of the paper "Learning to Highlight Audio by Watching Movies"☆15Oct 1, 2025Updated 10 months ago
- Machine learning speaker characteristics☆46Updated this week
- Signal processing and deep learning approaches to detect partial discharge faults in covered conductors. An entry for a Kaggle competitio…☆17Mar 30, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Modeling Harmonic Complexity using two models of Conditional Variational Autoencoders - MSc. Thesis☆10May 16, 2023Updated 3 years ago
- The electronic Holly Quran browser Elforkane☆11Nov 14, 2021Updated 4 years ago
- The implementation of "Instrument Separation of Symbolic Music by Explicitly Guided Diffusion Model"☆15Aug 16, 2022Updated 3 years ago
- code and demo of the ISMIR 2021 paper CollageNet☆12Jul 12, 2021Updated 5 years ago
- JAVA and Docker based solution to host all service components (Owner, Manufacturer, Rendezvous) defined in SDO protocol. Reuses binaries …☆11Apr 5, 2023Updated 3 years ago
- Implementation of "Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation"☆14Oct 31, 2024Updated last year
- Efficient solution to the initialization selection of DNNs in transfer learning using duality diagrams as a similarity measure framework.☆10Nov 17, 2020Updated 5 years ago
- Embedded Tajweed annotation for the Qur'an☆11Nov 30, 2025Updated 8 months ago
- 使用Sentencepiece对中文语料进行分词☆13Nov 30, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Efficient Methods for BEamforming Deconvolution☆20Oct 26, 2017Updated 8 years ago
- Real Acoustic Fields An Audio-Visual Room Acoustics Dataset and Benchmark☆64Aug 29, 2024Updated last year
- code for A Large-scale Dataset for Audio-Language Representation Learning☆14Sep 18, 2024Updated last year
- Learning an Interpretable End-to-End Network for Real-Time Acoustic Beamforming☆21Aug 20, 2024Updated last year
- code for "Automated and Intelligent Synthesis of Oxygen-Producing Catalysts from Martian Meteorites by Robotic AI-Chemist "☆12Jul 31, 2023Updated 3 years ago
- S3PRL for Speech Emotion Recognition (see s3prl > downstream)☆15Feb 28, 2026Updated 5 months ago
- ISMIR 2021: Curriculum Learning for Imbalanced Classification in Large Vocabulary Automatic Chord Recognition☆10Nov 8, 2021Updated 4 years ago
- Zenodo Developers Site☆18Apr 10, 2026Updated 3 months ago
- ☆19Aug 27, 2018Updated 7 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆17Sep 2, 2017Updated 8 years ago
- ☆11Jan 22, 2017Updated 9 years ago
- Tools to isolate speaker and transcribe unstructured audio clips☆11Dec 4, 2022Updated 3 years ago
- This repo contains the code to reproduce the paper: "Enriched Music Representations with Multiple Cross-modal Contrastive Learning"☆15Jun 22, 2023Updated 3 years ago
- soundvista☆16Dec 31, 2025Updated 7 months ago
- Voice Framework☆18Jan 21, 2026Updated 6 months ago
- Keras implementation of SincNet (https://github.com/mravanelli/SincNet, https://arxiv.org/abs/1808.00158)☆12Aug 5, 2018Updated 8 years ago