Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
☆47Mar 3, 2026Updated 5 months ago
Alternatives and similar repositories for Demo-ICL
Users that are interested in Demo-ICL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A local AI assistant running on your device. It turns your files into actionable memory.☆58Mar 24, 2026Updated 5 months ago
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆31Apr 2, 2026Updated 4 months ago
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence☆87Jul 22, 2026Updated last month
- ☆23Apr 11, 2026Updated 4 months ago
- SpatialBench: Is Your Spatial Foundation Model an All-Round Player?☆126Jul 25, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆44Mar 27, 2026Updated 5 months ago
- The official implementation of “MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction”☆67Mar 20, 2026Updated 5 months ago
- Privacy-first AI memory layer - Signal for AI Memory. E2EE, local-first, works with Claude, Cursor, and any MCP-compatible AI.☆23Jun 12, 2026Updated 2 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- [ArXiv 26] The official repository of "ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors".☆42Mar 5, 2026Updated 5 months ago
- A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Lan…☆23May 25, 2026Updated 3 months ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆27Apr 13, 2026Updated 4 months ago
- FileGram: Grounding Agent Personalization in File-System Behavioral Traces☆67Apr 12, 2026Updated 4 months ago
- Data release for Step Differences in Instructional Video (CVPR24)☆15Jun 19, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The official implementation of "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation". (arXiv 2601.22153)☆325Aug 14, 2026Updated 2 weeks ago
- [ECCV 2026] An official implementation of "EndoCoT". Scaling endogenous Chain-of-Thought (CoT) reasoning in diffusion models for complex …☆45Aug 11, 2026Updated 2 weeks ago
- Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos☆72Sep 5, 2025Updated 11 months ago
- Code for our IJCV paper "HumanLiff: Layer-wise 3D Human Generation with Diffusion Model"☆55Apr 11, 2026Updated 4 months ago
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding☆364Aug 5, 2026Updated 3 weeks ago
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 7 months ago
- ☆35Aug 12, 2026Updated 2 weeks ago
- Official Implementation of "Kinema4D: Kinematic4D World Modeling for Spatiotemporal Embodied Simulation"☆83May 21, 2026Updated 3 months ago
- Official repository of paper "LOVE-R1: Advancing Long Video Understanding with Adaptive Zoom-in Mechanism via Multi-Step Reasoning"☆24Nov 1, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [CVPR2025] Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models☆21Apr 30, 2025Updated last year
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆168Jun 10, 2026Updated 2 months ago
- [ICLR 2025] MLLM for On-Demand Spatial-Temporal Understanding at Arbitrary Resolution☆329Jul 4, 2025Updated last year
- ConsistentNeRF Enhances Neural Radiance Fields with 3D Consistency for Sparse View Synthesis☆74Oct 12, 2023Updated 2 years ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆397Jun 20, 2026Updated 2 months ago
- [CVPR 2025] WildAvatar: Learning In-the-wild 3D Avatars from the Web☆131Mar 11, 2025Updated last year
- On-Device Domain Generalization☆48Nov 9, 2022Updated 3 years ago
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆181Aug 14, 2026Updated 2 weeks ago
- [CVPR 2025] EgoLife: Towards Egocentric Life Assistant☆458Mar 19, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- PhysX: Physical-Grounded 3D Asset Generation (NeurIPS 2025, Spotlight)☆389Dec 18, 2025Updated 8 months ago
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations☆30May 21, 2025Updated last year
- The official PyTorch implementation of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) '24 paper PREGO: online mistake detect…☆35Jun 9, 2025Updated last year
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning☆37Aug 28, 2025Updated last year
- [CVPR 2024] How to Configure Good In-Context Sequence for Visual Question Answering☆21May 28, 2025Updated last year
- ☆79May 4, 2025Updated last year
- [ NeurIPS 2024 D&B Track ] Implementation for "FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models"☆74Dec 27, 2024Updated last year