Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
☆47Mar 3, 2026Updated 6 months ago
Alternatives and similar repositories for Demo-ICL
Users that are interested in Demo-ICL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A local AI assistant running on your device. It turns your files into actionable memory.☆58Mar 24, 2026Updated 5 months ago
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆33Apr 2, 2026Updated 5 months ago
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence☆92Jul 22, 2026Updated last month
- ☆23Apr 11, 2026Updated 5 months ago
- SpatialBench: Is Your Spatial Foundation Model an All-Round Player?☆132Jul 25, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆46Mar 27, 2026Updated 5 months ago
- The official implementation of “MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction”☆71Sep 4, 2026Updated 2 weeks ago
- Privacy-first AI memory layer - Signal for AI Memory. E2EE, local-first, works with Claude, Cursor, and any MCP-compatible AI.☆23Jun 12, 2026Updated 3 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- [ArXiv 26] The official repository of "ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors".☆42Mar 5, 2026Updated 6 months ago
- Syphus: Automatic Instruction-Response Generation Pipeline☆14Dec 14, 2023Updated 2 years ago
- A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Lan…☆23May 25, 2026Updated 3 months ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆27Apr 13, 2026Updated 5 months ago
- Data release for Step Differences in Instructional Video (CVPR24)☆15Jun 19, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official implementation of "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation". (arXiv 2601.22153)☆334Sep 2, 2026Updated 2 weeks ago
- [ACL2025 Findings] Benchmarking Multihop Multimodal Internet Agents☆54Feb 27, 2025Updated last year
- [ECCV 2026] An official implementation of "EndoCoT". Scaling endogenous Chain-of-Thought (CoT) reasoning in diffusion models for complex …☆47Aug 11, 2026Updated last month
- (ICML 2026) Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search☆50Aug 7, 2026Updated last month
- Code for our IJCV paper "HumanLiff: Layer-wise 3D Human Generation with Diffusion Model"☆55Apr 11, 2026Updated 5 months ago
- Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos☆72Sep 5, 2025Updated last year
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 7 months ago
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding☆371Aug 5, 2026Updated last month
- ☆35Aug 12, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official repository of paper "LOVE-R1: Advancing Long Video Understanding with Adaptive Zoom-in Mechanism via Multi-Step Reasoning"☆24Nov 1, 2025Updated 10 months ago
- [CVPR2025] Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models☆21Apr 30, 2025Updated last year
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆169Jun 10, 2026Updated 3 months ago
- [ICLR 2025] MLLM for On-Demand Spatial-Temporal Understanding at Arbitrary Resolution☆329Jul 4, 2025Updated last year
- ConsistentNeRF Enhances Neural Radiance Fields with 3D Consistency for Sparse View Synthesis☆74Oct 12, 2023Updated 2 years ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆400Jun 20, 2026Updated 3 months ago
- [CVPR 2025] WildAvatar: Learning In-the-wild 3D Avatars from the Web☆131Mar 11, 2025Updated last year
- On-Device Domain Generalization☆48Nov 9, 2022Updated 3 years ago
- ☆40Mar 3, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025] EgoLife: Towards Egocentric Life Assistant☆462Mar 19, 2025Updated last year
- PhysX: Physical-Grounded 3D Asset Generation (NeurIPS 2025, Spotlight)☆392Dec 18, 2025Updated 9 months ago
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations☆30May 21, 2025Updated last year
- The official PyTorch implementation of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) '24 paper PREGO: online mistake detect…☆35Jun 9, 2025Updated last year
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning☆37Aug 28, 2025Updated last year
- [CVPR 2024] How to Configure Good In-Context Sequence for Visual Question Answering☆21May 28, 2025Updated last year
- ☆79May 4, 2025Updated last year