This repo holds the implementation of PAVE: Patching and Adapting Video Large Language Models (CVPR2025)
☆27Sep 6, 2025Updated 11 months ago
Alternatives and similar repositories for PAVE
Users that are interested in PAVE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"☆65Oct 9, 2025Updated 10 months ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 8 months ago
- Official Repository of RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning☆14Jul 9, 2025Updated last year
- Official code for M3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks☆21Jun 4, 2026Updated 2 months ago
- On Path to Multimodal Generalist: General-Level and General-Bench☆21Jul 11, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆17Oct 25, 2024Updated last year
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆22Feb 21, 2026Updated 5 months ago
- Code For Our Work: DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries [ECCV-2024]☆15Jul 11, 2024Updated 2 years ago
- Tensorflow 2.0.0 implementation of SPRT-TANDEM☆13Jun 21, 2022Updated 4 years ago
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 3 months ago
- Official Implementation for "ESCAPE: Encoding Super-keypoints for Category-Agnostic Pose Estimation", CVPR 2024.☆10Jun 17, 2024Updated 2 years ago
- ☆13Jun 26, 2022Updated 4 years ago
- ☆22Oct 25, 2024Updated last year
- VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.☆18May 31, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2023] Official code for "Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations"☆56Aug 8, 2023Updated 3 years ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- Official PyTorch implementation of the ICML 2023 paper "Adaptive IMLE for Few-shot Pretraining-free Generative Modelling "☆16Feb 13, 2025Updated last year
- Official implementaiton of RefAM: Attention Magnets for Zero-Shot Referral Segmentaiton☆16Feb 6, 2026Updated 6 months ago
- MAVERICS (Manually-vAlidated Vq^2a Examples fRom Image-Caption datasetS) is a suite of test-only benchmarks for visual question answering…☆13Feb 18, 2023Updated 3 years ago
- Official implementation of Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More☆25Feb 25, 2025Updated last year
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆50Oct 9, 2025Updated 10 months ago
- [CVPR 2025] RAP: Retrieval-Augmented Personalization☆87Jun 8, 2026Updated 2 months ago
- ☆21Oct 5, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Nov 16, 2020Updated 5 years ago
- ☆17Mar 14, 2024Updated 2 years ago
- code of cvpr26 paper Symphony☆17Apr 7, 2026Updated 4 months ago
- 👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)☆74Jan 20, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆18Jun 2, 2026Updated 2 months ago
- [CVPR 2025 & IJCV2026] Official PyTorch Code for "MMRL: Multi-Modal Representation Learning for Vision-Language Models" and its extension…☆116Apr 6, 2026Updated 4 months ago
- ☆21Jun 12, 2025Updated last year
- ☆35Feb 12, 2026Updated 6 months ago
- Official implementation of POODLE: Improving Few-shot Learning via Penalizing Out-of-Distribution Samples (NeurIPS 2021)☆14Aug 6, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding☆24Apr 9, 2026Updated 4 months ago
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- [ICLR 2026] Official code for paper: TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinf…☆27Jan 29, 2026Updated 6 months ago
- [2026 AAAI] Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation☆20Nov 8, 2025Updated 9 months ago
- [ICLR 2025] CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion☆56Jul 1, 2025Updated last year
- ☆34Sep 19, 2025Updated 10 months ago
- Official Implementation of Video-MA2MBA☆12Dec 3, 2024Updated last year