Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation: A framework for generating multimodal music by bridging different representations and enhancing generation with RAG.
☆28Jan 21, 2025Updated last year
Alternatives and similar repositories for MTM
Users that are interested in MTM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆32Nov 10, 2025Updated 8 months ago
- Video Background Music Generation Using Unpaired Audio-Visual Data☆33Oct 8, 2024Updated last year
- Source codes for the paper "Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning" (PDMER) which p…☆14Mar 24, 2025Updated last year
- Code accompayning ISMIR23 paper; TriAD: Capturing harmonics with 3D convolutions☆20Jul 19, 2024Updated 2 years ago
- official code for CVPR'24 paper Diff-BGM☆71Oct 12, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Codebase for ICLR' 23 paper- ''wav2tok: Deep Sequence Tokenizer for Audio Retrieval"☆36Jun 30, 2026Updated 3 weeks ago
- ☆20May 7, 2025Updated last year
- ☆33Dec 23, 2025Updated 7 months ago
- ☆12Mar 11, 2025Updated last year
- [ICLR 2025] Weighted-Reward Preference Optimization for Implicit Model Fusion☆14Mar 17, 2025Updated last year
- ScorePerformer: Expressive Piano Performance Rendering with Fine-Grained Control (ISMIR 2023)☆42Mar 10, 2025Updated last year
- [ICCV 2023] Video Background Music Generation: Dataset, Method and Evaluation☆78Mar 29, 2024Updated 2 years ago
- Controllable Group Choreography using Contrastive Diffusion (SIGGRAPH ASIA 2023)☆19Nov 25, 2025Updated 8 months ago
- ☆15Feb 6, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ISMIR 24 Supplementary Material☆14Oct 28, 2024Updated last year
- This is the accompanying repository to the paper - Automatic Estimation of Singing Voice Musical Dynamics☆16Oct 28, 2024Updated last year
- Repo for the IDESSAI 2024 course on modeling audio with discrete tokens.☆13Sep 13, 2024Updated last year
- MuChoMusic is a benchmark for evaluating music understanding in multimodal audio-language models.☆46Dec 3, 2024Updated last year
- The official implementation of work "AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward".☆19Mar 25, 2025Updated last year
- "Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification" ISMIR2025☆38Sep 11, 2025Updated 10 months ago
- Code for NeurIPS 2025 Paper “MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation”☆17May 21, 2026Updated 2 months ago
- Source code for the EMNLP 2025 paper “DM-Codec: Distilling Multimodal Representations for Speech Tokenization”☆57Jun 1, 2025Updated last year
- Audio Prompt Adapter: Unleashing music editing abilities for text-to-music with lightweight finetuning [ISMIR 2024]☆57Nov 10, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The Ai Music Generation Challenge 2022☆27May 29, 2024Updated 2 years ago
- Codes for ICASSP 2024 paper: BEAST: Online Joint Beat and Downbeat Tracking Based on Streaming Transformer. An online beat tracking syste…☆44Sep 11, 2024Updated last year
- [CVPR 2025] Repository of VidMuse☆140Jun 7, 2025Updated last year
- Repository for the paper "Combining audio control and style transfer using latent diffusion", accepted at ISMIR 2024☆67Feb 19, 2025Updated last year
- Official implementation of TISDiSS, a scalable framework for discriminative source separation.☆16Oct 19, 2025Updated 9 months ago
- Humos paper repository☆26Sep 6, 2025Updated 10 months ago
- Repo of the paper "Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model""☆15Jun 28, 2024Updated 2 years ago
- ☆16Mar 26, 2025Updated last year
- ☆24Dec 10, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Official repository for GraFPrint: an audio identification framework based on graph neural networks.☆41Sep 18, 2025Updated 10 months ago
- [ICMR 2025] Official Repository for The Paper, Let Network Decide What to Learn: Symbolic Music Understanding Model Based on Large-scale …☆19Aug 17, 2025Updated 11 months ago
- The official implementation of our paper "Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tu…☆109Jan 14, 2026Updated 6 months ago
- [Official Implementation] Acoustic Autoregressive Modeling 🔥☆74Aug 24, 2024Updated last year
- [ICASSP 2025] AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder☆14Mar 11, 2025Updated last year
- Official repo for BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning presented at ICASSP 2025☆16Nov 18, 2024Updated last year
- The official GitHub page for the survey paper "Foundation Models for Music: A Survey".☆224Sep 4, 2024Updated last year