Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
☆47Mar 3, 2026Updated 7 months ago
Alternatives and similar repositories for Demo-ICL
Users that are interested in Demo-ICL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A local AI assistant running on your device. It turns your files into actionable memory.☆58Mar 24, 2026Updated 6 months ago
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆33Apr 2, 2026Updated 6 months ago
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence☆93Jul 22, 2026Updated 2 months ago
- [NeurIPS 2026] SpatialBench: Is Your Spatial Foundation Model an All-Round Player?☆138Jul 25, 2026Updated 2 months ago
- ☆48Mar 27, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The official implementation of “MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction”☆71Sep 4, 2026Updated last month
- Privacy-first AI memory layer - Signal for AI Memory. E2EE, local-first, works with Claude, Cursor, and any MCP-compatible AI.☆23Jun 12, 2026Updated 3 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- [ArXiv 26] The official repository of "ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors".☆42Mar 5, 2026Updated 7 months ago
- Syphus: Automatic Instruction-Response Generation Pipeline☆14Dec 14, 2023Updated 2 years ago
- A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Lan…☆23May 25, 2026Updated 4 months ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆27Apr 13, 2026Updated 5 months ago
- Data release for Step Differences in Instructional Video (CVPR24)☆15Jun 19, 2024Updated 2 years ago
- The official implementation of "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation". (NeurIPS 2026)☆350Sep 25, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ECCV 2026] An official implementation of "EndoCoT". Scaling endogenous Chain-of-Thought (CoT) reasoning in diffusion models for complex …☆48Aug 11, 2026Updated last month
- (ICML 2026) Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search☆50Aug 7, 2026Updated 2 months ago
- Code for our IJCV paper "HumanLiff: Layer-wise 3D Human Generation with Diffusion Model"☆54Apr 11, 2026Updated 5 months ago
- Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos☆73Sep 5, 2025Updated last year
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding☆363Aug 5, 2026Updated 2 months ago
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 8 months ago
- ☆36Aug 12, 2026Updated last month
- Official repository of paper "LOVE-R1: Advancing Long Video Understanding with Adaptive Zoom-in Mechanism via Multi-Step Reasoning"☆25Nov 1, 2025Updated 11 months ago
- [CVPR2025] Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models☆21Apr 30, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆169Jun 10, 2026Updated 3 months ago
- Official Implementation of "Kinema4D: Kinematic4D World Modeling for Spatiotemporal Embodied Simulation"☆83May 21, 2026Updated 4 months ago
- [ICLR 2025] MLLM for On-Demand Spatial-Temporal Understanding at Arbitrary Resolution☆329Jul 4, 2025Updated last year
- ConsistentNeRF Enhances Neural Radiance Fields with 3D Consistency for Sparse View Synthesis☆74Oct 12, 2023Updated 2 years ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆403Jun 20, 2026Updated 3 months ago
- [CVPR 2025] WildAvatar: Learning In-the-wild 3D Avatars from the Web☆131Mar 11, 2025Updated last year
- On-Device Domain Generalization☆48Nov 9, 2022Updated 3 years ago
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆185Aug 14, 2026Updated last month
- ☆41Mar 3, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [CVPR 2025] EgoLife: Towards Egocentric Life Assistant☆469Mar 19, 2025Updated last year
- PhysX: Physical-Grounded 3D Asset Generation (NeurIPS 2025, Spotlight)☆392Dec 18, 2025Updated 9 months ago
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations☆30May 21, 2025Updated last year
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning☆37Aug 28, 2025Updated last year
- [CVPR 2024] How to Configure Good In-Context Sequence for Visual Question Answering☆21May 28, 2025Updated last year
- ☆79May 4, 2025Updated last year
- [ NeurIPS 2024 D&B Track ] Implementation for "FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models"☆75Dec 27, 2024Updated last year