Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
☆260May 5, 2025Updated last year
Alternatives and similar repositories for GUI-R1
Users that are interested in GUI-R1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [AAAI 2026] Code for "UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning"☆160Nov 24, 2025Updated 9 months ago
- (ACL 2025) 🔥🔥🔥Code for "Empowering Multimodal Large Language Models with Evol-Instruct"☆21May 15, 2025Updated last year
- Repository for the paper "InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners"☆67Dec 4, 2025Updated 9 months ago
- ☆22Jan 9, 2026Updated 8 months ago
- [NeurIPS'25] GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents☆416Apr 13, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models☆59Jan 19, 2024Updated 2 years ago
- QuantClaw is a plug-and-play task-type routing quantization plugin for OpenClaw.☆116Apr 27, 2026Updated 4 months ago
- Official Implementation of ARPO: End-to-End Policy Optimization for GUI Agents with Experience Replay☆163May 29, 2025Updated last year
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost☆122Jul 17, 2025Updated last year
- Code repo for "Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning"☆34Jul 25, 2025Updated last year
- NeurIPS'2022: Pluralistic Image Completion with Gaussian Mixture Models☆14Jan 28, 2023Updated 3 years ago
- [MM 2025] Towards Modality Generalization: A Benchmark and Prospective Analysis☆31May 22, 2025Updated last year
- Official code for DeepSound-V1☆12May 14, 2025Updated last year
- (NIPS 2025) OpenOmni: Official implementation of Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Align…☆143May 9, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Uni-OVSeg is a weakly supervised open-vocabulary segmentation framework that leverages unpaired mask-text pairs.☆54Jun 11, 2024Updated 2 years ago
- Awesome GUI Agent Paper List☆904Updated this week
- DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning…☆30Sep 7, 2025Updated last year
- [ICML2025] Official Code of From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection☆27Jun 27, 2025Updated last year
- [NeurIPS 2025]"Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning"☆110Oct 21, 2025Updated 10 months ago
- ☆18May 14, 2025Updated last year
- ☆27Sep 15, 2025Updated last year
- ☆17May 2, 2024Updated 2 years ago
- The model, data and code for the visual GUI Agent SeeClick☆493Jul 13, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [AAAI 2026] GUI-G²: Gaussian Reward Modeling for GUI Grounding☆312Apr 15, 2026Updated 5 months ago
- This repository contains the code for our ICML 2025 paper——LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection🎉☆27May 29, 2025Updated last year
- This is the official repository of the paper "Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Schedulin…☆15Jul 27, 2025Updated last year
- ☆31Sep 12, 2025Updated last year
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis☆178Jun 18, 2026Updated 3 months ago
- DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation☆97Feb 26, 2026Updated 6 months ago
- 💻 A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.☆1,217Aug 17, 2025Updated last year
- [CVPR 2026] ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps☆87Jul 26, 2026Updated last month
- [ACL 2026] Repository of IPBench☆23Apr 6, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆100Dec 23, 2025Updated 8 months ago
- A curated collection of resources, tools, and frameworks for developing GUI Agents.☆457Jul 28, 2026Updated last month
- [NeurIPS 2025] UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents☆61Nov 27, 2025Updated 9 months ago
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,166Aug 31, 2026Updated 2 weeks ago
- [ICML 2026] GUIEvalKit: Open-source Evaluation Toolkit for GUI Agents☆25Feb 26, 2026Updated 6 months ago
- [ACM MM 2026] MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments☆49Sep 4, 2026Updated 2 weeks ago
- Regularly Truncated M-estimators for Learning with Noisy Labels☆11Apr 24, 2024Updated 2 years ago