We introduce 'Thinking with Video', a new paradigm leveraging video generation for multimodal reasoning. Our VideoThinkBench shows that Sora-2 surpasses GPT5 by 10% on eyeballing puzzles and reaches 69% accuracy on MMMU.
☆320Aug 23, 2026Updated last week
Alternatives and similar repositories for Thinking-with-Video
Users that are interested in Thinking-with-Video are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- We propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervis…☆431Jul 22, 2026Updated last month
- Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning☆159Aug 21, 2026Updated last week
- A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical …☆61Sep 1, 2025Updated 11 months ago
- ☆145Jun 24, 2026Updated 2 months ago
- MOSS-VL is the core multimodal model series within the OpenMOSS ecosystem, dedicated to visual understanding.☆484Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICML 2026] Prism: Spectral-Aware Block-Sparse Attention☆27May 22, 2026Updated 3 months ago
- An open-source personal academic homepage template characterized by its user-friendly design and extensive scalability.☆36Oct 6, 2025Updated 10 months ago
- Thinking with Videos from Open-Source Priors. We reproduce chain-of-frames visual reasoning by fine-tuning open-source video models. Give…☆229Apr 13, 2026Updated 4 months ago
- [ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"☆193May 1, 2026Updated 3 months ago
- [CVPR 2026] TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models☆67Feb 21, 2026Updated 6 months ago
- Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.1983…☆100Mar 9, 2026Updated 5 months ago
- ☆25Jan 29, 2026Updated 7 months ago
- [ICML 2025] M-STAR (Multimodal Self-Evolving TrAining for Reasoning) Project. Diving into Self-Evolving Training for Multimodal Reasoning☆75Jul 13, 2025Updated last year
- Cambrian-S: Towards Spatial Supersensing in Video☆567Apr 3, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This is a collection of recent papers on reasoning in video generation models.☆164Updated this week
- Are Video Models Ready as Zero-shot Reasoners?☆88Nov 24, 2025Updated 9 months ago
- ☆28Jan 22, 2026Updated 7 months ago
- [ICML 2026] Official repo for "DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models"☆186Jan 4, 2026Updated 7 months ago
- [ICLR2026] Laser: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping☆68May 22, 2025Updated last year
- PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails☆22Jul 8, 2026Updated last month
- [CVPR 2026] Official repo for "VideoSSR: Video Self-Supervised Reinforcement Learning"☆48Nov 11, 2025Updated 9 months ago
- MOVA: Towards Scalable and Synchronized Video–Audio Generation☆1,105Updated this week
- The official code repository for the FullFront benchmark☆28May 16, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- 启智平台的 Agent 驾驶舱:Skill + CLI,一条命令直达。Agent cockpit for the Inspire ML platform: one command, every operation, straight from chat.☆200Updated this week
- ☆23Mar 2, 2026Updated 5 months ago
- https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT☆139Jan 30, 2026Updated 6 months ago
- Native Multimodal Models are World Learners☆1,547Dec 30, 2025Updated 7 months ago
- [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactiv…☆938Updated this week
- ☆23Nov 21, 2025Updated 9 months ago
- ☆19Oct 12, 2025Updated 10 months ago
- Code repository for the ICML 2026 paper "Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation".☆24Jun 14, 2026Updated 2 months ago
- [ACL' 25] The official code repository for PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.☆94Feb 15, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"☆83Apr 3, 2026Updated 4 months ago
- ✈️ [ICCV 2025] Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints☆81Jul 10, 2025Updated last year
- Official Repository for paper "HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding" [ACL 2026]☆101May 8, 2026Updated 3 months ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆27Apr 13, 2026Updated 4 months ago
- [ICLR26] Understanding VS. Generation: Navigating Optimization Dilemma in Multimodal Models☆28May 6, 2026Updated 3 months ago
- Official code of "RoboOmni: Proactive Robot Manipulation in Omni-modal Context"☆119Mar 28, 2026Updated 5 months ago
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆56Oct 9, 2025Updated 10 months ago