MutiModel paper reading (Visual, Audio)
☆22Nov 24, 2025Updated 10 months ago
Alternatives and similar repositories for MLLM-paper-reading
Users that are interested in MLLM-paper-reading are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆51Apr 5, 2026Updated 5 months ago
- Official repository for “Duo-Tok: Dual-Track Semantic Music Tokenizer for Vocal–Accompaniment Generation.”☆32Nov 26, 2025Updated 10 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music☆128Jan 7, 2026Updated 8 months ago
- Code for ChordSync, a conformer-based audio-to-chord synchroniser☆14Oct 17, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆16Jan 9, 2026Updated 8 months ago
- ☆16Aug 10, 2025Updated last year
- [INTERSPEECH 2026] This is the official implementation for εar-VAE model including inference and evaluation parts, more details coming so…☆98Aug 22, 2026Updated last month
- trying to reproduce suno v3☆36Jan 29, 2025Updated last year
- ☆35Sep 6, 2025Updated last year
- 2025年深圳大学办公区校园网新版登录脚本。2025 Shenzhen University Office Area Campus Network New Version Login Script☆11Jan 17, 2025Updated last year
- ☆60Apr 30, 2026Updated 4 months ago
- Implementation of Acoustic BPE (Shen et al., 2024), extended for RVQ-based Neural Audio Codecs☆77Dec 3, 2025Updated 9 months ago
- Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-Step High-Fidelity Audio Generation☆149Mar 8, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Unconditional music synthesis using a diffusion model in the STFT domain☆12May 31, 2022Updated 4 years ago
- Constant-Q harmonic coefficients (CQHCs), a timbre feature designed for music signals.☆29Sep 13, 2025Updated last year
- Self-training LLaVA for medical☆17Nov 3, 2024Updated last year
- State-of-the-art pretrained music models for training, evaluation, inference☆190Aug 29, 2026Updated 3 weeks ago
- 一个基于Stable diffusion 1.5的中国山水画风的LoRA模型与其训练集和训练方法,并提供其扩展的与svd模型共同构建的文生视频工作流,并利用开源模型Anytext生成带有特定中国书法 的山水画。☆14Aug 16, 2024Updated 2 years ago
- Code from blog 'Searching by Music: Leveraging Vector Search for Music Information Retrieval'☆17Nov 16, 2023Updated 2 years ago
- ☆56Jul 13, 2025Updated last year
- ☆21Apr 24, 2025Updated last year
- ISMIR 24 Supplementary Material☆14Oct 28, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- code for Model-Guided Multi-Contrast Deep Unfolding Network for MRI Super-resolution Reconstruction☆16Oct 23, 2023Updated 2 years ago
- 一个用于在命令行环境下登陆深大校园网的客户端, 适用于 srun 认证系统☆19Aug 2, 2026Updated last month
- Official repository of Myna: Masking-Based Contrastive Learning of Musical Representations☆18Mar 31, 2025Updated last year
- ☆15Sep 26, 2022Updated 4 years ago
- A PyTorch implementation of the Modified Discrete Cosine Transform (MDCT) and its inverse for audio processing.☆33Sep 16, 2026Updated last week
- Parameter search for MIDI alignment☆17Sep 25, 2015Updated 11 years ago
- Official implementation of "AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and sta…☆51Nov 11, 2025Updated 10 months ago
- A graph-based deep MIDI music generator. Official repository of the paper "Graph-based Polyphonic Multitrack Music Generation".☆21Oct 13, 2023Updated 2 years ago
- Reference implementation of DecDTW in PyTorch (ICLR 2023)☆24May 29, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code and demo for paper: Zhao et al., "Q&A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement," IJCAI 202…☆21May 2, 2024Updated 2 years ago
- [ICLR 2025] Official PyTorch implementation of our paper for general continual learning "Advancing Prompt-Based Methods for Replay-Indepe…☆18Dec 21, 2025Updated 9 months ago
- This repository contains the official "LLM-as-a-Judge" evaluation scripts for the MixAssist project, as detailed in our paper, "MixAssist…☆19Jul 24, 2026Updated 2 months ago
- Code for the paper "Toward Fully Self-Supervised Multi-Pitch Estimation".☆25Sep 27, 2025Updated last year
- Python code to reproduce the experiments presented in the paper Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Ite…☆12Nov 13, 2020Updated 5 years ago
- ☆92Feb 24, 2026Updated 7 months ago
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆129Jun 4, 2025Updated last year