MutiModel paper reading (Visual, Audio)
☆22Nov 24, 2025Updated 8 months ago
Alternatives and similar repositories for MLLM-paper-reading
Users that are interested in MLLM-paper-reading are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆50Apr 5, 2026Updated 4 months ago
- Official repository for “Duo-Tok: Dual-Track Semantic Music Tokenizer for Vocal–Accompaniment Generation.”☆32Nov 26, 2025Updated 8 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music☆128Jan 7, 2026Updated 7 months ago
- Code for ChordSync, a conformer-based audio-to-chord synchroniser☆14Oct 17, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Jan 9, 2026Updated 7 months ago
- ☆16Aug 10, 2025Updated last year
- [INTERSPEECH 2026] This is the official implementation for εar-VAE model including inference and evaluation parts, more details coming so…☆91Feb 13, 2026Updated 6 months ago
- trying to reproduce suno v3☆35Jan 29, 2025Updated last year
- ☆35Sep 6, 2025Updated 11 months ago
- 2025年深圳大学办公区校园网新版登录脚本。2025 Shenzhen University Office Area Campus Network New Version Login Script☆11Jan 17, 2025Updated last year
- Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-Step High-Fidelity Audio Generation☆146Mar 8, 2026Updated 5 months ago
- Implementation of Acoustic BPE (Shen et al., 2024), extended for RVQ-based Neural Audio Codecs☆77Dec 3, 2025Updated 8 months ago
- Coupled Dictionary Learning based Multi-contrast MRI Reconstruction☆13Jun 16, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Unconditional music synthesis using a diffusion model in the STFT domain☆12May 31, 2022Updated 4 years ago
- Constant-Q harmonic coefficients (CQHCs), a timbre feature designed for music signals.☆29Sep 13, 2025Updated 11 months ago
- Self-training LLaVA for medical☆16Nov 3, 2024Updated last year
- State-of-the-art pretrained music models for training, evaluation, inference☆186Jul 30, 2026Updated 2 weeks ago
- 一个基于Stable diffusion 1.5的中国山水画风的LoRA模型与其训练集和训练方法,并提供其扩展的与svd模型共同构建的文 生视频工作流,并利用开源模型Anytext生成带有特定中国书法的山水画。☆13Aug 16, 2024Updated 2 years ago
- Diff-SFCT: A Diffusion Model with Spatial-Frequency Cross Transformer for Medical Image Segmentation☆10Apr 15, 2024Updated 2 years ago
- Code from blog 'Searching by Music: Leveraging Vector Search for Music Information Retrieval'☆16Nov 16, 2023Updated 2 years ago
- ☆56Jul 13, 2025Updated last year
- ☆21Apr 24, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ISMIR 24 Supplementary Material☆14Oct 28, 2024Updated last year
- code for Model-Guided Multi-Contrast Deep Unfolding Network for MRI Super-resolution Reconstruction☆16Oct 23, 2023Updated 2 years ago
- 一个用于在命令行环境下登陆深大校园网的客户端, 适用于 srun 认证系统☆18Aug 2, 2026Updated 2 weeks ago
- ☆15Sep 26, 2022Updated 3 years ago
- Official repository of Myna: Masking-Based Contrastive Learning of Musical Representations☆17Mar 31, 2025Updated last year
- Official implementation of "AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and sta…☆50Nov 11, 2025Updated 9 months ago
- A PyTorch implementation of the Modified Discrete Cosine Transform (MDCT) and its inverse for audio processing.☆33Dec 17, 2024Updated last year
- Python code to reproduce the experiments presented in the paper Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Ite…☆12Nov 13, 2020Updated 5 years ago
- Parameter search for MIDI alignment☆17Sep 25, 2015Updated 10 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A graph-based deep MIDI music generator. Official repository of the paper "Graph-based Polyphonic Multitrack Music Generation".☆21Oct 13, 2023Updated 2 years ago
- Reference implementation of DecDTW in PyTorch (ICLR 2023)☆24May 29, 2023Updated 3 years ago
- Code and demo for paper: Zhao et al., "Q&A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement," IJCAI 202…☆21May 2, 2024Updated 2 years ago
- IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)☆16Nov 15, 2024Updated last year
- This repository contains the official "LLM-as-a-Judge" evaluation scripts for the MixAssist project, as detailed in our paper, "MixAssist…☆16Jul 24, 2026Updated 3 weeks ago
- [ICLR 2025] Official PyTorch implementation of our paper for general continual learning "Advancing Prompt-Based Methods for Replay-Indepe…☆18Dec 21, 2025Updated 7 months ago
- Code for the paper "Toward Fully Self-Supervised Multi-Pitch Estimation".☆25Sep 27, 2025Updated 10 months ago