The official code and data for paper "VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI"
☆20Mar 25, 2025Updated last year
Alternatives and similar repositories for VidEgoThink
Users that are interested in VidEgoThink are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR'24 Highlight] The official code and data for paper "EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Lan…☆67Mar 25, 2025Updated last year
- ☆16Sep 25, 2025Updated 11 months ago
- Pytorch implementation for Egoinstructor at CVPR 2024☆28Dec 1, 2024Updated last year
- Code for NeurIPS 2022 Datasets and Benchmarks paper - EgoTaskQA: Understanding Human Tasks in Egocentric Videos.☆47Apr 17, 2023Updated 3 years ago
- (ICCV2025) Official repository of paper "ViSpeak: Visual Instruction Feedback in Streaming Videos"☆54Jul 1, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"☆20Jan 18, 2026Updated 7 months ago
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆168Jun 10, 2026Updated 2 months ago
- Agentic Keyframe Search for Video Question Answering☆18Jun 30, 2026Updated 2 months ago
- NLPBench: Evaluating NLP-Related Problem-solving Ability in Large Language Models☆10Oct 27, 2023Updated 2 years ago
- [ICLR 26] Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow Perspective☆22Mar 6, 2026Updated 5 months ago
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆23Feb 21, 2026Updated 6 months ago
- ☆14May 13, 2025Updated last year
- [CVPR'25] 🌟🌟 EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering☆52Jun 19, 2025Updated last year
- Code for MonoJSG: Joint Semantic and Geometric Cost Volume for Monocular 3D Object Detection☆32Oct 10, 2022Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆13Jan 14, 2022Updated 4 years ago
- Official Pytorch implementation of EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens [ICML2024].☆31Jun 15, 2024Updated 2 years ago
- Code and Dataset for the CVPRW Paper "Where did I leave my keys? — Episodic-Memory-Based Question Answering on Egocentric Videos"☆32Aug 28, 2023Updated 3 years ago
- ☆118Dec 30, 2024Updated last year
- Home Action Genome: Cooperative Contrastive Action Understanding☆22Nov 8, 2021Updated 4 years ago
- FGLA: Fast Generation-Based Gradient Leakage Attacks against Highly Compressed Gradients☆15Mar 17, 2026Updated 5 months ago
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"☆27Feb 22, 2026Updated 6 months ago
- For Ego4D VQ3D Task☆22Jan 9, 2024Updated 2 years ago
- ☆29Jul 25, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆18Jul 15, 2025Updated last year
- [ICLR'26] ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning☆16Feb 28, 2026Updated 6 months ago
- Röttger et al. (2025): "MSTS: A Multimodal Safety Test Suite for Vision-Language Models"☆20Mar 31, 2025Updated last year
- Code repository for MMUGL: Multi-modal Graph Learning over UMLS Knowledge Graphs☆11Dec 7, 2023Updated 2 years ago
- Energy-based Out-of-distribution Detection☆17Dec 23, 2020Updated 5 years ago
- [CVPR 2025] EgoLife: Towards Egocentric Life Assistant☆459Mar 19, 2025Updated last year
- ☆17Dec 30, 2022Updated 3 years ago
- A pytorch reimplementation of KL-Loss (CVPR'2019)☆15Oct 15, 2023Updated 2 years ago
- ☆41Sep 9, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Repository of paper: Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models☆37Sep 19, 2023Updated 2 years ago
- [ICML 2023] "Unleashing Mask: Explore the Intrinsic Out-of-Distribution Detection Capability"☆18Jul 7, 2023Updated 3 years ago
- This is the official repository for the paper "Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness…☆17Dec 18, 2025Updated 8 months ago
- [ACCV 2022 Oral] SymmNeRF: Learning to Explore Symmetry Prior for Single-View View Synthesis☆14Mar 14, 2024Updated 2 years ago
- 此代码用于RoboMaster AI Challenge 2020的平面仿真☆10May 10, 2020Updated 6 years ago
- Source code for our paper: "LoGU: Long-form Generation with Uncertainty Expressions".☆19May 27, 2025Updated last year
- [ICML 2024] Learning Reward for Robot Skills Using Large Language Models via Self-Alignment☆19Aug 22, 2024Updated 2 years ago