AudioBERT π’ : Audio Knowledge Augmented Language Model (ICASSP 2025)
β40Feb 1, 2025Updated last year
Alternatives and similar repositories for AudioBERT
Users that are interested in AudioBERT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of "The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers (ECCV 2024)ββ26Jan 15, 2025Updated last year
- Codes for "Learning bounds for risk-sensitive learning," NeurIPS 2020 (or see arXiv 2006.08138)β11Oct 15, 2020Updated 5 years ago
- Official Implementation of "Prefix tuning for Automated Audio Captioning(ICASSP 2023)"β30Dec 6, 2023Updated 2 years ago
- Research Papers on Efficient Neural Fields from EffL Groupβ16Apr 21, 2025Updated last year
- Meta-Learning Sparse Implicit Neural Representations (NeurIPS 2021)β64Oct 29, 2021Updated 4 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- In progress.β69Mar 26, 2024Updated 2 years ago
- UTAUTAI(Unrestricted Tune Automated Technology Artificial Interigence)β17Oct 27, 2023Updated 2 years ago
- β19Feb 2, 2023Updated 3 years ago
- An All-in-One Speech, Sound, Music Codec with Single Nested Codebookβ28Oct 11, 2025Updated 9 months ago
- A lightweight audio codec based on a single quantizerβ72Aug 15, 2025Updated 11 months ago
- β18May 4, 2025Updated last year
- Baseline for DCASE 2024 Task 9: "Language-Queried Audio Source Separation"β26Mar 27, 2024Updated 2 years ago
- β12Mar 11, 2025Updated last year
- Unofficial implementation JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation(https://arxiv.org/abs/2310.1β¦β32Jan 19, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The open source code for SimpleSpeech seriesβ147Oct 8, 2024Updated last year
- High-quality Text-to-Audio Generation with Efficient Diffusion Transformerβ333Dec 17, 2025Updated 7 months ago
- SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANsβ16Jul 19, 2023Updated 3 years ago
- [DEPRECIATED] [PyTorch 2.0] [638M] [85.33% acc] Full-attention multi-instrumental music transformer for supervised music generation, optiβ¦β33Nov 23, 2023Updated 2 years ago
- [ICASSP 2026]Official code for "Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum"β27Jan 22, 2026Updated 6 months ago
- Code for the paper: GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilitiesβ153Dec 5, 2024Updated last year
- Codebase and project page for EDMSoundβ35Nov 20, 2023Updated 2 years ago
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformersβ126Mar 20, 2025Updated last year
- The source code for the paper CrossSinger (asru2023)β18Oct 12, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The source code for the paper XiaoiceSing2 (interspeech2023)β49Jan 15, 2024Updated 2 years ago
- Codebase for the paper 'EncodecMAE: Leveraging neural codecs for universal audio representation learning'β101Jul 24, 2024Updated 2 years ago
- Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval (TTMR++) [ICASSP24]β43Oct 7, 2024Updated last year
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modellingβ100Nov 9, 2024Updated last year
- SOTA Piano Transformer model trained on 4.2GB of Solo Piano MIDI musicβ28Nov 9, 2023Updated 2 years ago
- [InterSpeech 24] FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filterβ98Jul 4, 2024Updated 2 years ago
- SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.β121Jan 28, 2026Updated 5 months ago
- Inference codebase for "Cacophony: An Improved Contrastive Audio-Text Model". Preprint: https://arxiv.org/abs/2402.06986β49Jan 19, 2026Updated 6 months ago
- Musical Word Embedding for Music Tagging and Retrieval [IEEE TASLP]β29Apr 23, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- MuChoMusic is a benchmark for evaluating music understanding in multimodal audio-language models.β46Dec 3, 2024Updated last year
- Official Implementation of "Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity (ICML 2024)"β44Aug 28, 2024Updated last year
- DEX-TTS: Diffusion-based EXpressive TTS with Style Modeling on Time Variabilityβ108Jan 17, 2025Updated last year
- [ACL 2026 Main] FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pre-trainingβ36Apr 20, 2026Updated 3 months ago
- Variable Bitrate Residual Vector Quantization for Audio Codingβ54May 1, 2025Updated last year
- Official repository of Wavehax vocoderβ75Dec 20, 2025Updated 7 months ago
- [NAACL 2025] WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matchingβ133Apr 8, 2026Updated 3 months ago