[TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
β162Jun 10, 2026Updated 3 weeks ago
Alternatives and similar repositories for Ego-R1
Users that are interested in Ego-R1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π§ VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)β343Feb 8, 2026Updated 4 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMMβ20May 22, 2025Updated last year
- [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?β94Jul 13, 2025Updated 11 months ago
- [AAAI 2025] Grounded Multi-Hop VideoQA in Long-Form Egocentric Videosβ38May 27, 2025Updated last year
- [CVPR 2025] EgoLife: Towards Egocentric Life Assistantβ442Mar 19, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Modelβ92Nov 27, 2025Updated 7 months ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmarkβ25Apr 13, 2026Updated 2 months ago
- β100Jun 23, 2025Updated last year
- MR. Video: MapReduce is the Principle for Long Video Understandingβ31Jun 18, 2026Updated 2 weeks ago
- β16Sep 25, 2025Updated 9 months ago
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)β725Sep 24, 2025Updated 9 months ago
- TStar is a unified temporal search framework for long-form video question answeringβ97Mar 23, 2026Updated 3 months ago
- [CVPR2025] Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editingβ26Aug 23, 2025Updated 10 months ago
- Code for "Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning [EMNLP 2025 Findings]"β18Aug 27, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β34Feb 12, 2026Updated 4 months ago
- Streaming Video Instruction Tuningβ76Feb 25, 2026Updated 4 months ago
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluationβ20Jun 2, 2025Updated last year
- TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoningβ115Dec 24, 2025Updated 6 months ago
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceβ84May 28, 2026Updated last month
- The official implement of "Grounded Chain-of-Thought for Multimodal Large Language Models"β22Jul 21, 2025Updated 11 months ago
- Video-R1: Reinforcing Video Reasoning in MLLMs [π₯the first paper to explore R1 for video]β878Dec 14, 2025Updated 6 months ago
- \infty-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidationβ21Feb 14, 2025Updated last year
- β70Feb 27, 2026Updated 4 months ago
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [ICML 2025] LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Modelsβ17Nov 4, 2025Updated 7 months ago
- MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignmentβ35Jul 1, 2024Updated 2 years ago
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"β50Oct 9, 2025Updated 8 months ago
- [NeurIPS 2024] Official code for HourVideo: 1-Hour Video Language Understandingβ146Jul 12, 2025Updated 11 months ago
- This repository contains the implementation of the paper: "ChatCam: Empowering Camera Control through Conversational AI", NeurIPS 2024.β22Nov 15, 2024Updated last year
- β158Oct 31, 2024Updated last year
- TTRV: Test-Time Reinforcement Learning for VisionβLanguage Models (CVPR 2026)β45Mar 8, 2026Updated 3 months ago
- [ACL2025 Oral & Award] Evaluate Image/Video Generation like Humans - Fast, Explainable, Flexibleβ127Aug 10, 2025Updated 10 months ago
- Code implementation of paper "MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval (AAAI2025)"β26Feb 2, 2025Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- β20May 11, 2025Updated last year
- Diffusion Powers Video Tokenizer for Comprehension and Generation (CVPR 2025)β88Feb 27, 2025Updated last year
- Long Context Transfer from Language to Visionβ406Mar 18, 2025Updated last year
- The code for "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by VIdeo SpatioTemporal Augmentation" [CVPR2025]β20Feb 27, 2025Updated last year
- PhysGame Benchmark for Physical Commonsense Evaluation in Gameplay Videosβ49Jul 3, 2025Updated last year
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learningβ55Jul 23, 2025Updated 11 months ago
- [π IJCV 2025 & ACCV 2024 Best Paper Honorable Mention] Official pytorch implementation of the paper "High-Quality Visually-Guided Sound β¦β33Mar 30, 2026Updated 3 months ago