This repo holds the implementation of PAVE: Patching and Adapting Video Large Language Models (CVPR2025)
☆27Sep 6, 2025Updated 10 months ago
Alternatives and similar repositories for PAVE
Users that are interested in PAVE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"☆65Oct 9, 2025Updated 9 months ago
- Video feature extraction pipeline that supports diverse models including I3D, SlowFast, EgoVLP, and CLIP.☆13Apr 20, 2024Updated 2 years ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 7 months ago
- Official Repository of RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning☆14Jul 9, 2025Updated last year
- Official code for M3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks☆21Jun 4, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Template job submissions using GPUs in CHTC☆54Updated this week
- On Path to Multimodal Generalist: General-Level and General-Bench☆21Jul 11, 2025Updated last year
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆21Feb 21, 2026Updated 5 months ago
- Code For Our Work: DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries [ECCV-2024]☆15Jul 11, 2024Updated 2 years ago
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 3 months ago
- ☆13Jun 26, 2022Updated 4 years ago
- [CVPR2025] We present SleeperMark, a novel framework designed to embed resilient watermarks into T2I diffusion models☆39May 26, 2025Updated last year
- ☆22Oct 25, 2024Updated last year
- VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.☆18May 31, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆10Mar 23, 2025Updated last year
- [CVPR 2023] Official code for "Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations"☆56Aug 8, 2023Updated 2 years ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- MAVERICS (Manually-vAlidated Vq^2a Examples fRom Image-Caption datasetS) is a suite of test-only benchmarks for visual question answering…☆13Feb 18, 2023Updated 3 years ago
- Code release for RICA^2: Rubric-Informed, Calibrated Assessment of Actions (ECCV 2024)☆15Nov 9, 2025Updated 8 months ago
- Official implementation of Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More☆25Feb 25, 2025Updated last year
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆50Oct 9, 2025Updated 9 months ago
- [CVPR 2025] RAP: Retrieval-Augmented Personalization☆87Jun 8, 2026Updated last month
- ☆12Nov 16, 2020Updated 5 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- fork from https://github.com/jwyang/faster-rcnn.pytorch☆10Aug 6, 2018Updated 7 years ago
- ☆17Mar 14, 2024Updated 2 years ago
- code of cvpr26 paper Symphony☆17Apr 7, 2026Updated 3 months ago
- 👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)☆74Jan 20, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆18Jun 2, 2026Updated last month
- [CVPR 2025 & IJCV2026] Official PyTorch Code for "MMRL: Multi-Modal Representation Learning for Vision-Language Models" and its extension…☆116Apr 6, 2026Updated 3 months ago
- ☆34Feb 12, 2026Updated 5 months ago
- ☆21Jun 12, 2025Updated last year
- OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding☆22Apr 9, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Dataset of measurements from a low-cost single-photon camera used in our CVPR 2024 paper "Towards 3D Vision with Low-Cost Single-Photon C…☆15Nov 24, 2025Updated 8 months ago
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- [ICLR 2026] Official code for paper: TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinf…☆27Jan 29, 2026Updated 5 months ago
- [2026 AAAI] Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation☆20Nov 8, 2025Updated 8 months ago
- [ICLR 2025] CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion☆56Jul 1, 2025Updated last year
- Official Implementation of Video-MA2MBA☆12Dec 3, 2024Updated last year
- Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆26Apr 5, 2026Updated 3 months ago