[CVPR 2026] HiconAgent: History Context-aware Policy Optimization for GUI Agents
☆31Mar 9, 2026Updated 4 months ago
Alternatives and similar repositories for CVPR26-HiconAgent
Users that are interested in CVPR26-HiconAgent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICCV 2025 Highlight] Less is More: Empowering GUI Agent with Context-Aware Simplification☆48Mar 12, 2026Updated 4 months ago
- [ACM MM 2025] PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning☆18Jun 6, 2026Updated last month
- [ACL 2026 main] PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records☆21Apr 11, 2026Updated 3 months ago
- [TIP 2026] UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries☆34May 7, 2026Updated 2 months ago
- [CVPR 2025] LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant☆29Dec 2, 2025Updated 7 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ACM MM 2025] Official repository of "EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoni…☆30May 6, 2026Updated 2 months ago
- [CVPR 2026] ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation☆46May 8, 2026Updated 2 months ago
- [CVPR 2026] Official Implementation for Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Effi…☆25Updated this week
- [NeurIPS 2025] CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification☆185Jun 17, 2026Updated last month
- [AAAI 2026 Oral] SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation☆70Apr 5, 2026Updated 3 months ago
- Support finetuning GLM4v with zero2☆16Jun 29, 2024Updated 2 years ago
- [MM‘25] GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting☆19Nov 24, 2025Updated 7 months ago
- [CVPR 2026] Implementation of HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding☆20Apr 30, 2026Updated 2 months ago
- [NeurIPS 2024] Official Implementation for Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks☆102Jun 17, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A curated list of large VLM-based VLA models for robotic manipulation.☆426Apr 3, 2026Updated 3 months ago
- Under construction☆13Jan 15, 2025Updated last year
- [NeurIPS 2024] MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models☆85Dec 27, 2025Updated 6 months ago
- A simple visual test-time scaling method for GUI agent grounding☆26Dec 7, 2025Updated 7 months ago
- EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments☆28May 15, 2025Updated last year
- [arxiv: 2512.19673] Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies☆60Feb 6, 2026Updated 5 months ago
- ☆20May 20, 2025Updated last year
- This repo contains evaluation code for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?"☆31Dec 23, 2024Updated last year
- Python codes for mathematical modeling.☆13Sep 5, 2021Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- CaMML:Context-Aware MultiModal Learner for Large Models (ACL 2024 SAC Award)☆15May 21, 2025Updated last year
- A Chinese-focused PyTorch framework for exploring Attention Residuals in Qwen3-style causal LMs, with baseline, Block AttnRes, Full AttnR…☆19May 3, 2026Updated 2 months ago
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.☆10May 16, 2024Updated 2 years ago
- MMICL, a state-of-the-art VLM with the in context learning ability from ICL, PKU☆49Jul 13, 2025Updated last year
- [ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"☆69Apr 3, 2026Updated 3 months ago
- F1: A Vision Language Action Model Bridging Understanding and Generation to Actions☆200Jan 2, 2026Updated 6 months ago
- ☆34Oct 27, 2024Updated last year
- This is the official implementation of 2025 CVPR paper "EmoEdit: Evoking Emotions through Image Manipulation".☆40Apr 27, 2026Updated 2 months ago
- Self-Supervised Dataset Distillation for Transfer Learning☆19Apr 10, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [TCSVT23] Official code for "SPT: Spatial Pyramid Transformer for Image Captioning".☆10Aug 14, 2024Updated last year
- [ICLR 2025] Official Implementation for 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3…☆18Apr 7, 2026Updated 3 months ago
- Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents☆252May 5, 2025Updated last year
- [ACL 2025] GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent☆68May 28, 2025Updated last year
- Arbitrary Entropy Policy Optimization: Entropy Is Controllable in Reinforcement Fine-tuning☆17Jan 19, 2026Updated 6 months ago
- ☆14Nov 14, 2023Updated 2 years ago
- ORES: Open-vocabulary Responsible Visual Synthesis☆14Dec 12, 2023Updated 2 years ago