A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
☆39Sep 22, 2025Updated 11 months ago
Alternatives and similar repositories for minimal_video_pairs
Users that are interested in minimal_video_pairs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models’ under…☆62Aug 18, 2025Updated last year
- Official code for MotionBench (CVPR 2025)☆79Mar 3, 2025Updated last year
- [CVPR 2025 Highlight] Official implementation of HySAC, a hyperbolic safety-aware vision-language model for safer multimodal retrieval an…☆31Apr 8, 2025Updated last year
- Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"☆41Aug 12, 2026Updated 2 weeks ago
- An unofficial implementation: Dual Refinement Underwater Object Detection Network☆11May 11, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16Nov 23, 2023Updated 2 years ago
- [ICCV 2023] Official PyTorch implementation of the paper "DiffTAD: Temporal Action Detection with Proposal Denoising Diffusion"☆37Mar 30, 2023Updated 3 years ago
- Library that provides metrics to assess representation quality☆30Feb 5, 2025Updated last year
- [CVPR 2025 Oral] VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection☆142Jul 28, 2025Updated last year
- [CVPR 2025] Hyperbolic Category Discovery☆33Nov 7, 2025Updated 9 months ago
- ☆34Sep 19, 2025Updated 11 months ago
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models☆28Nov 29, 2025Updated 9 months ago
- Awesome Vision-Language Compositionality, a comprehensive curation of research papers in literature.☆41Feb 13, 2025Updated last year
- ☆38Feb 6, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for the paper: "Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation"☆60Sep 22, 2025Updated 11 months ago
- Code and data for ImageCoDe, a contextual vison-and-language benchmark☆42Mar 1, 2024Updated 2 years ago
- LiveClin is a live benchmark designed for the faithful replication of clinical practice☆17Feb 27, 2026Updated 6 months ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 8 months ago
- PyTorch code and models for VJEPA2 self-supervised learning from video.☆4,551Mar 23, 2026Updated 5 months ago
- EARL: Editing with Autoregression and RL☆43Nov 21, 2025Updated 9 months ago
- Official repo for arxiv paper "Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion I…☆17Nov 8, 2024Updated last year
- [ECCV 2026] VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs☆18Feb 3, 2026Updated 6 months ago
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols☆20Nov 19, 2025Updated 9 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆78Apr 9, 2026Updated 4 months ago
- [Neurocomputing] Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation☆26Dec 21, 2025Updated 8 months ago
- Official Code Repo for the paper "Learning to Play Atari in a World of Tokens" accepted at ICML, 2024☆11Jun 6, 2024Updated 2 years ago
- "Large Language Models for Disease Diagnosis: A Scoping Review" (npj Artificial Intelligence 2025)☆18Mar 28, 2026Updated 5 months ago
- ☆20Jan 26, 2025Updated last year
- Repository for the CVPR23 paper Re^2TAL☆13Nov 21, 2025Updated 9 months ago
- The Source Code for WebCompass☆25May 2, 2026Updated 3 months ago
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- code for CoRL2025 "LaDiWM: A Latent Diffusion-based World Model for Predictive Manipulation"☆58Nov 30, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official code for Temporal Straightening for Latent Planning☆107Apr 27, 2026Updated 4 months ago
- [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs☆62Feb 2, 2026Updated 6 months ago
- InvTorch: Memory-Efficient Invertible Functions☆17Oct 31, 2024Updated last year
- Official code and data from DexWM ("World Models Can Leverage Human Videos for Dexterous Manipulation").☆96Jun 23, 2026Updated 2 months ago
- [ICLR 2026] MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning☆45Jan 14, 2026Updated 7 months ago
- Learning about objects and their properties by interacting with them☆12Oct 21, 2020Updated 5 years ago
- [CVPR 2024] Official repository of ST_GT☆10Sep 15, 2024Updated last year