Official code for MotionBench (CVPR 2025)
☆76Mar 3, 2025Updated last year
Alternatives and similar repositories for MotionBench
Users that are interested in MotionBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Accepted By The 39th Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track☆25Nov 17, 2025Updated 8 months ago
- [ICLR 2026] "VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?", Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu, L…☆41Jan 30, 2026Updated 5 months ago
- ☆13Apr 13, 2026Updated 3 months ago
- [ICLR 2026] MotionSight's official code implementation.☆48Apr 24, 2026Updated 2 months ago
- F-16 is a powerful video large language model (LLM) that perceives high-frame-rate videos, which is developed by the Department of Electr…☆40Jul 3, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos☆72Sep 5, 2025Updated 10 months ago
- Pixels, Patterns, but no Poetry: To See the World like Humans☆18Aug 11, 2025Updated 11 months ago
- [ICLR 2026] Official implementation of the paper "Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs"☆24Mar 3, 2026Updated 4 months ago
- ☆11Aug 4, 2024Updated last year
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs☆39Sep 22, 2025Updated 10 months ago
- Extending context length of visual language models☆12Dec 18, 2024Updated last year
- ☆41Nov 8, 2024Updated last year
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆42Apr 28, 2026Updated 2 months ago
- CoMA: Compositional Human Motion Generation with Multi-modal Agents☆16Jul 31, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT☆137Jan 30, 2026Updated 5 months ago
- ☆18Apr 9, 2026Updated 3 months ago
- [ICCV 2025] Implementation of the paper "Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs"☆81Oct 25, 2025Updated 8 months ago
- [ICML 2026] Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions☆48Jun 29, 2026Updated 3 weeks ago
- TStar is a unified temporal search framework for long-form video question answering☆97Mar 23, 2026Updated 4 months ago
- [CVPR 2025] PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models☆54Jun 12, 2025Updated last year
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆18Jul 15, 2025Updated last year
- ☆28Aug 9, 2025Updated 11 months ago
- [CVPR 2025 Oral] VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection☆140Jul 28, 2025Updated 11 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation☆20Jun 2, 2025Updated last year
- ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer☆42Jan 29, 2026Updated 5 months ago
- ☆30Sep 4, 2025Updated 10 months ago
- (3DV 2026) Pytorch implementation of “InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos”☆28Mar 16, 2026Updated 4 months ago
- (ICCV2025) Official repository of paper "ViSpeak: Visual Instruction Feedback in Streaming Videos"☆53Jul 1, 2025Updated last year
- TVBench: Redesigning Video-Language Evaluation☆15Jun 9, 2025Updated last year
- Scaling Motion Generation Model with Million-Level Human Motions (ICML 2025)☆69May 14, 2025Updated last year
- ☆13May 17, 2025Updated last year
- A benchmark that focuses on the sampling dilemma in long-video tasks. Through well-designed tasks, it evaluates the sampling efficiency o…☆28Aug 7, 2025Updated 11 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICLR 2025] Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation☆53Mar 13, 2025Updated last year
- ☆16Updated this week
- [Arxiv 2024] MotionCLR: Motion Generation and Training-free Editing via Understanding Attention Mechanisms☆16Dec 1, 2024Updated last year
- ☆15Jun 2, 2025Updated last year
- The official repository for paper "FlexSelect: Flexible Token Selection for Efficient Long Video Understanding".☆31Sep 19, 2025Updated 10 months ago
- ☆29Jul 23, 2025Updated last year
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)☆726Sep 24, 2025Updated 9 months ago