[CVPR 2025] EgoLife: Towards Egocentric Life Assistant
☆459Mar 19, 2025Updated last year
Alternatives and similar repositories for EgoLife
Users that are interested in EgoLife are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆168Jun 10, 2026Updated 2 months ago
- A local AI assistant running on your device. It turns your files into actionable memory.☆58Mar 24, 2026Updated 5 months ago
- [CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe☆166Mar 30, 2026Updated 5 months ago
- A framework that allows you to apply Sparse AutoEncoder on any models☆53Jul 11, 2025Updated last year
- [CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning☆108Jul 29, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆397Jun 20, 2026Updated 2 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated last month
- Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos☆72Sep 5, 2025Updated 11 months ago
- [CVPR'25] 🌟🌟 EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering☆52Jun 19, 2025Updated last year
- [ICCV 2025] Auto Interpretation Pipeline and many other functionalities for Multimodal SAE Analysis.☆200Sep 26, 2025Updated 11 months ago
- Privacy-first AI memory layer - Signal for AI Memory. E2EE, local-first, works with Claude, Cursor, and any MCP-compatible AI.☆23Jun 12, 2026Updated 2 months ago
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆182Aug 14, 2026Updated 2 weeks ago
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆112Mar 14, 2025Updated last year
- Holistic Evaluation of Multimodal LLMs on Spatial Intelligence☆123Jul 1, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model☆95Nov 27, 2025Updated 9 months ago
- Official implementation of EgoThinker at NIPS 2025☆29Nov 25, 2025Updated 9 months ago
- Cambrian-S: Towards Spatial Supersensing in Video☆568Apr 3, 2026Updated 4 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆20Apr 2, 2025Updated last year
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)☆729Sep 24, 2025Updated 11 months ago
- Official repo and evaluation implementation of VSI-Bench☆738Aug 5, 2025Updated last year
- ☆1,443Feb 12, 2026Updated 6 months ago
- [CVPR 2025] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?☆161Jul 24, 2025Updated last year
- [ICCV 2025] The official implementation for EgoM2P: Egocentric Multimodal Multitask Pretraining.☆42Jun 15, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]☆888Dec 14, 2025Updated 8 months ago
- Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition☆47Mar 3, 2026Updated 5 months ago
- One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks☆4,390Updated this week
- ☆79May 4, 2025Updated last year
- ☆4,717Jun 15, 2026Updated 2 months ago
- Ego4d dataset repository. Download the dataset, visualize, extract features & example usage of the dataset☆638Jul 25, 2026Updated last month
- Benchmarking and Analyzing Generative Data for Visual Recognition☆27Jul 25, 2023Updated 3 years ago
- The official implementation of "Compositional Generative Model of Unbounded 4D Cities". (TPAMI 2026)☆151Dec 6, 2025Updated 8 months ago
- Fully Open Framework for Democratized Multimodal Training☆1,195Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- NEO Series: Native Vision-Language Models from First Principles☆888Jul 27, 2026Updated last month
- The official repo of TeleEgo - A Benchmark for Egocentric AI Assistants.☆66Updated this week
- Long Context Transfer from Language to Vision☆412Mar 18, 2025Updated last year
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence☆87Jul 22, 2026Updated last month
- [ICLR 2026] MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning☆45Jan 14, 2026Updated 7 months ago
- T* keyframe search for long-form video understanding (CVPR 2025) + LV-Haystack temporal search benchmark code☆97Aug 23, 2026Updated last week
- Benchmark and training code for MindCube: spatial mental modeling in vision-language models from limited views.☆168Aug 23, 2026Updated last week