[ICCV 2025] This repo is the official implementation of "Music Grounding by Short Video"
☆26Sep 9, 2025Updated 11 months ago
Alternatives and similar repositories for MGSV
Users that are interested in MGSV are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official source code of our AAAI25 paper "D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matchin…☆10Feb 9, 2025Updated last year
- ☆10Nov 27, 2024Updated last year
- [ICCV 2023] Simple Baselines for Interactive Video Retrieval with Questions and Answers☆20Apr 16, 2024Updated 2 years ago
- [CVPR 2024] TeachCLIP for Text-to-Video Retrieval☆42May 7, 2025Updated last year
- [ICLR 2026] Empowering Small VLMs to Think with Dynamic Memorization and Exploration☆18Mar 18, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR 2026] GIR-Bench: Versatile Benchmark for Generating Images with Reasoning☆36Jan 27, 2026Updated 7 months ago
- ☆15Dec 20, 2024Updated last year
- 🕵️ ArXiv Agent v1.0 - Your Intelligent Research Assistant☆26Dec 29, 2025Updated 8 months ago
- [NeurIPS 2019] Drill-down: Interactive Retrieval of Complex Scenes using Natural Language Queries☆12Apr 15, 2022Updated 4 years ago
- Official Repository for "Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality" (ECCV 2024)☆16Oct 29, 2024Updated last year
- Stable-V2A: Synthesis of Synchronized Sound Effect with Temporal and Semantic Controls☆18May 27, 2025Updated last year
- calvis: Chest, wAist and peLVIS circumference from 3D human Body meshes for Deep Learning.☆14Jul 16, 2026Updated last month
- Official implementation of "Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound". IEEE TASLP 20…☆19Feb 27, 2026Updated 6 months ago
- Streaming Video Instruction Tuning☆87Feb 25, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- My notes for cmu15445 2022☆14Feb 8, 2023Updated 3 years ago
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM☆27Updated this week
- ☆12Feb 22, 2024Updated 2 years ago
- [2025 CVPR] Towards Open-Vocabulary Audio-Visual Event Localization☆47Mar 7, 2025Updated last year
- Boosting Multi-view Stereo with Late Cost Aggregation☆13Jan 24, 2024Updated 2 years ago
- Code implementation of the paper 'ExpertAF: Expert Actionable Feedback from Video'☆19Sep 30, 2025Updated 11 months ago
- [ICML2026] OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models☆28May 21, 2026Updated 3 months ago
- LLaVA-Next for STVG☆21Dec 5, 2025Updated 8 months ago
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ACL 20] Probing Linguistic Features of Sentence-level Representations in Neural Relation Extraction☆13Apr 21, 2020Updated 6 years ago
- Pytorch implementation of the paper 'Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Super…☆19Jan 19, 2024Updated 2 years ago
- [MMM‘24 Oral]CT-MVSNet: Efficient Multi-View Stereo with Cross-scale Transformer☆18Apr 18, 2024Updated 2 years ago
- [EMNLP'2023 Findings] MoqaGPT, for zero-shot multimodal question answering with LLMs☆13Dec 28, 2024Updated last year
- [TIP25] Code for "Text-Video Retrieval with Global-Local Semantic Consistent Learning"☆16May 12, 2025Updated last year
- The official implementation of the paper "Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study"☆18May 8, 2026Updated 3 months ago
- [CVPR 2022] The code for our paper 《Object-aware Video-language Pre-training for Retrieval》☆61May 25, 2022Updated 4 years ago
- Efficient and User-Friendly Time Series Analysis Library for PyOpenTS with pytorch compatibility.☆16Aug 14, 2023Updated 3 years ago
- [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection☆30Sep 26, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆11Aug 20, 2025Updated last year
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆35Feb 11, 2026Updated 6 months ago
- [ICME 2025] ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo☆17May 26, 2026Updated 3 months ago
- Code for A Dual Domain Multi-exposure Image Fusion Network Based on the Spatial-frequency Integration.☆12Jul 25, 2024Updated 2 years ago
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]☆15Jun 1, 2026Updated 3 months ago
- [TGRS 2023] Point Label Meets Remote Sensing Change Detection: A Consistency-Aligned Regional Growth Network☆15Jan 5, 2024Updated 2 years ago
- F-16 is a powerful video large language model (LLM) that perceives high-frame-rate videos, which is developed by the Department of Electr…☆40Jul 3, 2025Updated last year