[CVPR 2026] HiconAgent: History Context-aware Policy Optimization for GUI Agents
☆33Mar 9, 2026Updated 6 months ago
Alternatives and similar repositories for CVPR26-HiconAgent
Users that are interested in CVPR26-HiconAgent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACM MM 2025] PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning☆19Jun 6, 2026Updated 3 months ago
- [ACL 2026 main] PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records☆28Apr 11, 2026Updated 5 months ago
- [TIP 2026] UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries☆34May 7, 2026Updated 4 months ago
- [CVPR 2025] LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant☆29Dec 2, 2025Updated 9 months ago
- [ACM MM 2025] Official repository of "EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoni…☆32May 6, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [CVPR 2026] ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation☆50May 8, 2026Updated 4 months ago
- [AAAI 2026] H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Reffnement for Robotic Manipulation☆32Nov 28, 2025Updated 9 months ago
- [CVPR 2026] Official Implementation for Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Effi…☆34Jul 16, 2026Updated 2 months ago
- [NeurIPS 2025] CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification☆191Jun 17, 2026Updated 3 months ago
- [AAAI 2026 Oral] SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation☆71Apr 5, 2026Updated 5 months ago
- [NeurIPS 2024] Official Implementation for Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks☆103Jun 17, 2025Updated last year
- [NeurIPS 2024] MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models☆87Dec 27, 2025Updated 8 months ago
- [MM‘25] GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting☆21Aug 2, 2026Updated last month
- AAAI 2026 Oral☆21Dec 23, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official Implementation for Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimiza…☆72Apr 14, 2026Updated 5 months ago
- [CVPR 2026] Implementation of HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding☆25Aug 2, 2026Updated last month
- A curated list of large VLM-based VLA models for robotic manipulation.☆447Aug 27, 2026Updated 3 weeks ago
- EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments☆28May 15, 2025Updated last year
- [EMNLP 2026 Main] Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies☆60Feb 6, 2026Updated 7 months ago
- This is the official repository of the paper "Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Schedulin…☆15Jul 27, 2025Updated last year
- [ACL 2026 Oral] Official implementation of LaMI: Augmenting Large Language Models via Late Multi-Image Fusion☆20Jul 4, 2026Updated 2 months ago
- ☆20May 20, 2025Updated last year
- DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation☆97Feb 26, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This repo contains evaluation code for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?"☆31Dec 23, 2024Updated last year
- GPT Demo with hybrid distributed training☆10Dec 1, 2022Updated 3 years ago
- ☆18Mar 9, 2023Updated 3 years ago
- A Chinese-focused PyTorch framework for exploring Attention Residuals in Qwen3-style causal LMs, with baseline, Block AttnRes, Full AttnR…☆21May 3, 2026Updated 4 months ago
- CaMML:Context-Aware MultiModal Learner for Large Models (ACL 2024 SAC Award)☆15May 21, 2025Updated last year
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.☆10May 16, 2024Updated 2 years ago
- MMICL, a state-of-the-art VLM with the in context learning ability from ICL, PKU☆49Jul 13, 2025Updated last year
- [ICCV 2025] Official implementation of "What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?"☆21Aug 7, 2025Updated last year
- ☆66Sep 6, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- F1: A Vision Language Action Model Bridging Understanding and Generation to Actions☆199Jan 2, 2026Updated 8 months ago
- Predicting the onset of Alzheimer's Disease using MRI & PET scans.☆11Dec 18, 2018Updated 7 years ago
- ☆35Oct 27, 2024Updated last year
- [BMVC 2022] Information Theoretic Representation Distillation☆19Oct 6, 2023Updated 2 years ago
- [TCSVT23] Official code for "SPT: Spatial Pyramid Transformer for Image Captioning".☆10Aug 14, 2024Updated 2 years ago
- 电子科技大学(UESTC)编译技术实验☆15Dec 2, 2022Updated 3 years ago
- This is the official implementation of 2025 CVPR paper "EmoEdit: Evoking Emotions through Image Manipulation".☆42Apr 27, 2026Updated 4 months ago