☆20Apr 24, 2024Updated 2 years ago
Alternatives and similar repositories for seeclick-crawler
Users that are interested in seeclick-crawler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GUICourse: From General Vision Langauge Models to Versatile GUI Agents☆144Mar 1, 2026Updated 6 months ago
- The model, data and code for the visual GUI Agent SeeClick☆491Jul 13, 2025Updated last year
- ☆12Aug 8, 2024Updated 2 years ago
- ☆18Nov 1, 2024Updated last year
- ☆11Jul 6, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR 2024] Trajectory-as-Exemplar Prompting with Memory for Computer Control☆71Jan 7, 2026Updated 8 months ago
- Official implementation for "You Only Look at Screens: Multimodal Chain-of-Action Agents" (Findings of ACL 2024)☆263Jul 16, 2024Updated 2 years ago
- [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents☆320Aug 24, 2026Updated last month
- [ICLR 2025] A trinity of environments, tools, and benchmarks for general virtual agents☆233Jun 16, 2025Updated last year
- A Universal Platform for Training and Evaluation of Mobile Interaction☆64Sep 24, 2025Updated last year
- [ICCV 2025] GUIOdyssey is a comprehensive dataset for training and evaluating cross-app navigation agents. GUIOdyssey consists of 8,834 e…☆162Jan 3, 2026Updated 8 months ago
- [EMNLP 2022] The baseline code for META-GUI dataset☆16Jul 9, 2024Updated 2 years ago
- ☆34Updated this week
- GUI Grounding for Professional High-Resolution Computer Use☆398Jun 17, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆13Dec 9, 2024Updated last year
- This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception o…☆29Jul 9, 2025Updated last year
- Code for LaMPP: Language Models as Probabilistic Priors for Perception and Action☆37Apr 3, 2023Updated 3 years ago
- ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World☆26Jun 17, 2025Updated last year
- ☆62Sep 9, 2026Updated 2 weeks ago
- Consists of ~500k human annotations on the RICO dataset identifying various icons based on their shapes and semantics, and associations b…☆36Jun 27, 2024Updated 2 years ago
- Retrieved Sequence Augmentation for Protein Representation Learning☆51Nov 1, 2023Updated 2 years ago
- ☆19Nov 3, 2025Updated 10 months ago
- A pre labelled dataset for ui element / layout detection☆68Jun 15, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICML'24] TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks☆33Sep 20, 2024Updated 2 years ago
- [ICML2025] Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction☆393Mar 7, 2025Updated last year
- GPT-4V in Wonderland: LMMs as Smartphone Agents☆134Jul 17, 2024Updated 2 years ago
- Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments (EMNLP'2024)☆38Dec 29, 2024Updated last year
- (ICLR 2025) The Official Code Repository for GUI-World.☆70Aug 2, 2026Updated last month
- VisualWebArena is a benchmark for multimodal agents.☆487Nov 9, 2024Updated last year
- Trending projects & awesome papers about data-centric llm studies.☆41May 20, 2025Updated last year
- [CVPR 2024] Code and datasets for 'Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos'☆14Jun 16, 2024Updated 2 years ago
- Multimodal Open-O1 (MO1) is designed to enhance the accuracy of inference models by utilizing a novel prompt-based approach. This tool wo…☆27Sep 25, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The Screen Annotation dataset consists of pairs of mobile screenshots and their annotations. The annotations are in text format, and desc…☆94Mar 7, 2024Updated 2 years ago
- Official Repo for "Why Settle for One? Text-to-ImageSet Generation and Evaluation"☆22Oct 1, 2025Updated 11 months ago
- Urban Generative Intelligence (UGI): A Foundational Platform for Embodied Agent and Future City☆12Dec 17, 2023Updated 2 years ago
- ☆18Jul 31, 2026Updated last month
- Code for Research Project TLDR☆26Jul 28, 2025Updated last year
- Official repo for paper DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.☆397Feb 22, 2025Updated last year
- ☆24Jun 13, 2023Updated 3 years ago