Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
☆40Mar 3, 2026Updated 4 months ago
Alternatives and similar repositories for Demo-ICL
Users that are interested in Demo-ICL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A local AI assistant running on your device. It turns your files into actionable memory.☆55Mar 24, 2026Updated 3 months ago
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence☆73Jun 26, 2026Updated 3 weeks ago
- ☆23Apr 11, 2026Updated 3 months ago
- SpatialBench: Is Your Spatial Foundation Model an All-Round Player?☆114May 28, 2026Updated last month
- ☆42Mar 27, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official implementation of “MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction”☆66Mar 20, 2026Updated 4 months ago
- Privacy-first AI memory layer - Signal for AI Memory. E2EE, local-first, works with Claude, Cursor, and any MCP-compatible AI.☆23Jun 12, 2026Updated last month
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- [ArXiv 26] The official repository of "ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors".☆41Mar 5, 2026Updated 4 months ago
- Syphus: Automatic Instruction-Response Generation Pipeline☆14Dec 14, 2023Updated 2 years ago
- A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Lan…☆23May 25, 2026Updated last month
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆25Apr 13, 2026Updated 3 months ago
- Data release for Step Differences in Instructional Video (CVPR24)☆15Jun 19, 2024Updated 2 years ago
- The official implementation of "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation". (arXiv 2601.22153)☆311Jul 13, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL2025 Findings] Benchmarking Multihop Multimodal Internet Agents☆54Feb 27, 2025Updated last year
- [ECCV 2026] An official implementation of "EndoCoT". Scaling endogenous Chain-of-Thought (CoT) reasoning in diffusion models for complex …☆43Jun 26, 2026Updated 3 weeks ago
- (ICML 2026) Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search☆47May 1, 2026Updated 2 months ago
- Code for our IJCV paper "HumanLiff: Layer-wise 3D Human Generation with Diffusion Model"☆55Apr 11, 2026Updated 3 months ago
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding☆369May 24, 2026Updated last month
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 5 months ago
- ☆34Mar 14, 2026Updated 4 months ago
- Official repository of paper "LOVE-R1: Advancing Long Video Understanding with Adaptive Zoom-in Mechanism via Multi-Step Reasoning"☆24Nov 1, 2025Updated 8 months ago
- Official Implementation of "Kinema4D: Kinematic4D World Modeling for Spatiotemporal Embodied Simulation"☆78May 21, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆165Jun 10, 2026Updated last month
- [CVPR2025] Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models☆21Apr 30, 2025Updated last year
- [ICLR 2025] MLLM for On-Demand Spatial-Temporal Understanding at Arbitrary Resolution☆329Jul 4, 2025Updated last year
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆385Jun 20, 2026Updated last month
- [CVPR 2025] WildAvatar: Learning In-the-wild 3D Avatars from the Web☆130Mar 11, 2025Updated last year
- On-Device Domain Generalization☆47Nov 9, 2022Updated 3 years ago
- PhysX: Physical-Grounded 3D Asset Generation (NeurIPS 2025, Spotlight)☆380Dec 18, 2025Updated 7 months ago
- [CVPR 2025] EgoLife: Towards Egocentric Life Assistant☆447Mar 19, 2025Updated last year
- ☆40Mar 3, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆180Mar 23, 2025Updated last year
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations☆30May 21, 2025Updated last year
- [CVPR 2024] How to Configure Good In-Context Sequence for Visual Question Answering☆21May 28, 2025Updated last year
- ☆79May 4, 2025Updated last year
- The dataset repo of "CLCIFAR: CIFAR-Derived Benchmark Datasets with Human Annotated Complementary Labels" paper☆17May 11, 2026Updated 2 months ago
- [ NeurIPS 2024 D&B Track ] Implementation for "FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models"☆74Dec 27, 2024Updated last year
- 4D Panoptic Scene Graph Generation (NeurIPS'23 Spotlight)☆122Mar 13, 2025Updated last year