This repo holds the implementation of PAVE: Patching and Adapting Video Large Language Models (CVPR2025)
☆28Sep 6, 2025Updated last year
Alternatives and similar repositories for PAVE
Users that are interested in PAVE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"☆65Oct 9, 2025Updated 10 months ago
- ☆18Dec 23, 2022Updated 3 years ago
- Video feature extraction pipeline that supports diverse models including I3D, SlowFast, EgoVLP, and CLIP.☆13Apr 20, 2024Updated 2 years ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 8 months ago
- Official Repository of RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning☆15Jul 9, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official code for M3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks☆21Jun 4, 2026Updated 3 months ago
- On Path to Multimodal Generalist: General-Level and General-Bench☆22Jul 11, 2025Updated last year
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆23Feb 21, 2026Updated 6 months ago
- Code For Our Work: DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries [ECCV-2024]☆15Jul 11, 2024Updated 2 years ago
- In search of effective and efficient Pipeline for Distillating Knowledge in Convolutional Neural Networks☆13Apr 12, 2020Updated 6 years ago
- Tensorflow 2.0.0 implementation of SPRT-TANDEM☆14Jun 21, 2022Updated 4 years ago
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 4 months ago
- Official Implementation for "ESCAPE: Encoding Super-keypoints for Category-Agnostic Pose Estimation", CVPR 2024.☆10Jun 17, 2024Updated 2 years ago
- ☆13Jun 26, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [CVPR2025] We present SleeperMark, a novel framework designed to embed resilient watermarks into T2I diffusion models☆39May 26, 2025Updated last year
- ☆22Oct 25, 2024Updated last year
- VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.☆18May 31, 2026Updated 3 months ago
- [CVPR 2023] Official code for "Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations"☆56Aug 8, 2023Updated 3 years ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- Official PyTorch implementation of the ICML 2023 paper "Adaptive IMLE for Few-shot Pretraining-free Generative Modelling "☆16Aug 17, 2026Updated 2 weeks ago
- Official implementaiton of RefAM: Attention Magnets for Zero-Shot Referral Segmentaiton☆16Feb 6, 2026Updated 7 months ago
- MAVERICS (Manually-vAlidated Vq^2a Examples fRom Image-Caption datasetS) is a suite of test-only benchmarks for visual question answering…☆13Feb 18, 2023Updated 3 years ago
- Code release for RICA^2: Rubric-Informed, Calibrated Assessment of Actions (ECCV 2024)☆15Nov 9, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆56Oct 9, 2025Updated 10 months ago
- [CVPR 2025] RAP: Retrieval-Augmented Personalization☆86Jun 8, 2026Updated 2 months ago
- fork from https://github.com/jwyang/faster-rcnn.pytorch☆10Aug 6, 2018Updated 8 years ago
- ☆18Mar 14, 2024Updated 2 years ago
- code of cvpr26 paper Symphony☆17Apr 7, 2026Updated 5 months ago
- 👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)☆74Jan 20, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆20Jun 2, 2026Updated 3 months ago
- Code for AAAI 2024 paper "GCNext: Towards the Unity of Graph Convolutions for Human Motion Prediction"☆18Jan 16, 2025Updated last year
- [CVPR 2025 & IJCV2026] Official PyTorch Code for "MMRL: Multi-Modal Representation Learning for Vision-Language Models" and its extension…☆115Apr 6, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆23Jun 12, 2025Updated last year
- ☆35Feb 12, 2026Updated 6 months ago
- Official implementation of POODLE: Improving Few-shot Learning via Penalizing Out-of-Distribution Samples (NeurIPS 2021)☆14Aug 6, 2022Updated 4 years ago
- ☆18Nov 8, 2023Updated 2 years ago
- Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing☆63Dec 17, 2024Updated last year
- OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding☆25Apr 9, 2026Updated 4 months ago
- Implimentation of paper RAP☆16Feb 23, 2026Updated 6 months ago