AudioBERT π’ : Audio Knowledge Augmented Language Model (ICASSP 2025)
β40Feb 1, 2025Updated last year
Alternatives and similar repositories for AudioBERT
Users that are interested in AudioBERT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- UTAUTAI(Unrestricted Tune Automated Technology Artificial Interigence)β17Oct 27, 2023Updated 2 years ago
- β19Feb 2, 2023Updated 3 years ago
- An All-in-One Speech, Sound, Music Codec with Single Nested Codebookβ28Oct 11, 2025Updated 11 months ago
- A lightweight audio codec based on a single quantizerβ72Aug 15, 2025Updated last year
- β18May 4, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Baseline for DCASE 2024 Task 9: "Language-Queried Audio Source Separation"β26Mar 27, 2024Updated 2 years ago
- β12Mar 11, 2025Updated last year
- Unofficial implementation JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation(https://arxiv.org/abs/2310.1β¦β32Jan 19, 2024Updated 2 years ago
- The open source code for SimpleSpeech seriesβ146Oct 8, 2024Updated last year
- High-quality Text-to-Audio Generation with Efficient Diffusion Transformerβ332Dec 17, 2025Updated 9 months ago
- SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANsβ16Jul 19, 2023Updated 3 years ago
- [DEPRECIATED] [PyTorch 2.0] [638M] [85.33% acc] Full-attention multi-instrumental music transformer for supervised music generation, optiβ¦β33Nov 23, 2023Updated 2 years ago
- [ICASSP 2026]Official code for "Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum"β27Jan 22, 2026Updated 8 months ago
- Code for the paper: GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilitiesβ153Dec 5, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Codebase and project page for EDMSoundβ35Nov 20, 2023Updated 2 years ago
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformersβ127Mar 20, 2025Updated last year
- The source code for the paper CrossSinger (asru2023)β18Oct 12, 2023Updated 2 years ago
- The source code for the paper XiaoiceSing2 (interspeech2023)β49Jan 15, 2024Updated 2 years ago
- Codebase for the paper 'EncodecMAE: Leveraging neural codecs for universal audio representation learning'β101Jul 24, 2024Updated 2 years ago
- Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval (TTMR++) [ICASSP24]β42Oct 7, 2024Updated last year
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modellingβ100Nov 9, 2024Updated last year
- SOTA Piano Transformer model trained on 4.2GB of Solo Piano MIDI musicβ29Nov 9, 2023Updated 2 years ago
- [InterSpeech 24] FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filterβ99Jul 4, 2024Updated 2 years ago
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.β124Jan 28, 2026Updated 7 months ago
- Inference codebase for "Cacophony: An Improved Contrastive Audio-Text Model". Preprint: https://arxiv.org/abs/2402.06986β49Jan 19, 2026Updated 8 months ago
- Musical Word Embedding for Music Tagging and Retrieval [IEEE TASLP]β29Apr 23, 2024Updated 2 years ago
- MuChoMusic is a benchmark for evaluating music understanding in multimodal audio-language models.β46Dec 3, 2024Updated last year
- DEX-TTS: Diffusion-based EXpressive TTS with Style Modeling on Time Variabilityβ108Jan 17, 2025Updated last year
- [ACL 2026 Main] FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pre-trainingβ37Apr 20, 2026Updated 5 months ago
- Variable Bitrate Residual Vector Quantization for Audio Codingβ56May 1, 2025Updated last year
- Official repository of Wavehax vocoderβ78Dec 20, 2025Updated 9 months ago
- [NAACL 2025] WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matchingβ133Apr 8, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This repository provides the materials used in "Unsupervised Melody-to-Lyric Generation" by Yufei Tian, Anjali Narayan-Chen, Shereen Orabβ¦β11Jul 6, 2023Updated 3 years ago
- 5Hz Deep-Compression Speech VAE for AR-Diffusion and CALMsβ57Nov 19, 2025Updated 10 months ago
- β27Sep 10, 2025Updated last year
- [ACL 2025 Oral] Language-Codec: Reducing the Gaps Between Discrete Codec Representation and Speech Language Modelsβ208Jun 25, 2025Updated last year
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"β11Apr 10, 2025Updated last year
- The open source code of ALMTokenizer2: Towards Low bit-rate and Semantic-rich Audio Tokenizer with Flow-based Scalar Diffusion Transformeβ¦β46Sep 5, 2025Updated last year
- Implementation of "Audio xLSTMs: Learning Self-supervised audio representations with xLSTMs" in PyTorchβ20Aug 28, 2026Updated 3 weeks ago