MutiModel paper reading (Visual, Audio)
☆22Nov 24, 2025Updated 8 months ago
Alternatives and similar repositories for MLLM-paper-reading
Users that are interested in MLLM-paper-reading are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆50Apr 5, 2026Updated 3 months ago
- Official repository for “Duo-Tok: Dual-Track Semantic Music Tokenizer for Vocal–Accompaniment Generation.”☆32Nov 26, 2025Updated 8 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music☆27Jan 7, 2026Updated 6 months ago
- Code for ChordSync, a conformer-based audio-to-chord synchroniser☆14Oct 17, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆15Jan 9, 2026Updated 6 months ago
- ☆16Aug 10, 2025Updated 11 months ago
- [INTERSPEECH 2026] This is the official implementation for εar-VAE model including inference and evaluation parts, more details coming so…☆91Feb 13, 2026Updated 5 months ago
- trying to reproduce suno v3☆35Jan 29, 2025Updated last year
- ☆35Sep 6, 2025Updated 10 months ago
- 2025年深圳大学办公区校园网新版登录脚本。2025 Shenzhen University Office Area Campus Network New Version Login Script☆11Jan 17, 2025Updated last year
- Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-Step High-Fidelity Audio Generation☆146Mar 8, 2026Updated 4 months ago
- Implementation of Acoustic BPE (Shen et al., 2024), extended for RVQ-based Neural Audio Codecs☆76Dec 3, 2025Updated 7 months ago
- Unconditional music synthesis using a diffusion model in the STFT domain☆12May 31, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Constant-Q harmonic coefficients (CQHCs), a timbre feature designed for music signals.☆29Sep 13, 2025Updated 10 months ago
- State-of-the-art pretrained music models for training, evaluation, inference☆184Jan 20, 2026Updated 6 months ago
- 一个基于Stable diffusion 1.5的中国山水画风的LoRA模型与其训练集和训练方法,并提供其扩展的与svd模型共同构建的文生视频工作流,并利用开源模型Anytext生成带有特定中国书法的山水画。☆13Aug 16, 2024Updated last year
- Diff-SFCT: A Diffusion Model with Spatial-Frequency Cross Transformer for Medical Image Segmentation☆10Apr 15, 2024Updated 2 years ago
- Code from blog 'Searching by Music: Leveraging Vector Search for Music Information Retrieval'☆16Nov 16, 2023Updated 2 years ago
- ☆56Jul 13, 2025Updated last year
- ☆21Apr 24, 2025Updated last year
- ISMIR 24 Supplementary Material☆14Oct 28, 2024Updated last year
- 一个用于在命令行环境下登陆深大校园网的客户端, 适用于 srun 认证系统☆18Sep 25, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆15Sep 26, 2022Updated 3 years ago
- Official repository of Myna: Masking-Based Contrastive Learning of Musical Representations☆17Mar 31, 2025Updated last year
- Official implementation of "AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and sta…☆50Nov 11, 2025Updated 8 months ago
- A PyTorch implementation of the Modified Discrete Cosine Transform (MDCT) and its inverse for audio processing.☆33Dec 17, 2024Updated last year
- Python code to reproduce the experiments presented in the paper Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Ite…☆12Nov 13, 2020Updated 5 years ago
- Parameter search for MIDI alignment☆18Sep 25, 2015Updated 10 years ago
- A graph-based deep MIDI music generator. Official repository of the paper "Graph-based Polyphonic Multitrack Music Generation".☆21Oct 13, 2023Updated 2 years ago
- Reference implementation of DecDTW in PyTorch (ICLR 2023)☆24May 29, 2023Updated 3 years ago
- Code and demo for paper: Zhao et al., "Q&A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement," IJCAI 202…☆21May 2, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)☆17Nov 15, 2024Updated last year
- This repository contains the official "LLM-as-a-Judge" evaluation scripts for the MixAssist project, as detailed in our paper, "MixAssist…☆16Updated this week
- [ICLR 2025] Official PyTorch implementation of our paper for general continual learning "Advancing Prompt-Based Methods for Replay-Indepe…☆18Dec 21, 2025Updated 7 months ago
- Code for the paper "Toward Fully Self-Supervised Multi-Pitch Estimation".☆25Sep 27, 2025Updated 10 months ago
- ☆88Feb 24, 2026Updated 5 months ago
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆125Jun 4, 2025Updated last year
- official training and inference code of bitwise tokenizer☆71May 18, 2025Updated last year