CVPR25
☆28Jul 2, 2025Updated last year
Alternatives and similar repositories for MP-GUI
Users that are interested in MP-GUI are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A mobile GUI search engine using a vision-language model☆15May 5, 2025Updated last year
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- VisionDroid☆22Apr 2, 2024Updated 2 years ago
- ☆44Dec 8, 2025Updated 8 months ago
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding☆29Jun 10, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code repo for "Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding"☆31May 12, 2026Updated 2 months ago
- [NeurIPS 2025 Spotlight] Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning☆55Apr 16, 2026Updated 3 months ago
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models (ICLR2026)☆23Jun 24, 2026Updated last month
- ☆24Jul 8, 2023Updated 3 years ago
- 吴恩达大模型系列课程中文版,包括《Prompt Engineering》、《Building System》和《LangChain》☆12Jun 7, 2023Updated 3 years ago
- ☆19Sep 4, 2025Updated 11 months ago
- [ICCV 2025] GUIOdyssey is a comprehensive dataset for training and evaluating cross-app navigation agents. GUIOdyssey consists of 8,834 e…☆159Jan 3, 2026Updated 7 months ago
- ☆13Jul 30, 2026Updated last week
- ☆10Apr 22, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 非机动车头盔佩戴检测☆12Jan 20, 2026Updated 6 months ago
- Game UI Glitch Detection via Bug Understanding☆12Jul 31, 2021Updated 5 years ago
- ScreenQA dataset was introduced in the "ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots" paper. It contains ~86K …☆151Feb 7, 2025Updated last year
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement☆19Jul 4, 2026Updated last month
- Efficient Feature Extraction for High-resolution Video Frame Interpolation (BMVC 2022)☆14Aug 24, 2023Updated 2 years ago
- Under construction☆14Jan 15, 2025Updated last year
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference☆25May 21, 2026Updated 2 months ago
- ☆72Feb 27, 2026Updated 5 months ago
- This is the official repository for "Can GPTs Evaluate Graphic Design Based on Design Principles?".☆13Feb 10, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Owl Eyes: Spotting UI Display Issues via Visual Understanding☆12Jul 31, 2020Updated 6 years ago
- ☆12Aug 24, 2023Updated 2 years ago
- [TIP 2025] Advancing Zero-Shot Digital Human Quality Assessment through Text-Prompted Evaluation☆12Jul 8, 2023Updated 3 years ago
- This is the Reproducible Realisation of the AAAI25 paper "Look Inside for More: Internal Spatial Modality Perception for 3D Anomaly Detec…☆16Oct 5, 2025Updated 10 months ago
- [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents☆316Mar 11, 2026Updated 4 months ago
- ☆47Nov 8, 2024Updated last year
- ☆10Nov 9, 2023Updated 2 years ago
- ☆10Dec 3, 2024Updated last year
- This repository hosts the source code for the paper "ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Mo…☆16Dec 16, 2025Updated 7 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICML 2026] Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs☆33Apr 27, 2026Updated 3 months ago
- ☆11Oct 17, 2024Updated last year
- [ECCV 2024] FlexAttention for Efficient High-Resolution Vision-Language Models☆50Jan 8, 2025Updated last year
- [WACV 2026] ZonUI-3B — A lightweight, resolution-aware GUI grounding model trained with only 24K samples on a single RTX 4090.☆26Jan 2, 2026Updated 7 months ago
- [CVPR 2025 Oral] VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection☆141Jul 28, 2025Updated last year
- 字体识别☆12Apr 9, 2018Updated 8 years ago
- Emotiv SDK Community Edition☆13Oct 9, 2015Updated 10 years ago