[NeurIPS 2024 D&B] VideoGUI: A Benchmark for GUI Automation from Instructional Videos
β53Feb 22, 2026Updated 5 months ago
Alternatives and similar repositories for videogui
Users that are interested in videogui are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Repository of GUI Action Narratorβ13Apr 8, 2025Updated last year
- π» A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.β1,211Aug 17, 2025Updated 11 months ago
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation, ECCV 2024β22Feb 15, 2024Updated 2 years ago
- A Massive Multi-Discipline Lecture Understanding Benchmarkβ34Apr 20, 2026Updated 3 months ago
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluationβ20Jun 2, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICCV 2025] GUIOdyssey is a comprehensive dataset for training and evaluating cross-app navigation agents. GUIOdyssey consists of 8,834 eβ¦β159Jan 3, 2026Updated 7 months ago
- [CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.β1,891Apr 24, 2026Updated 3 months ago
- [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agentsβ316Mar 11, 2026Updated 5 months ago
- Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"β68Oct 19, 2024Updated last year
- VisualWebArena is a benchmark for multimodal agents.β485Nov 9, 2024Updated last year
- Data Journalist Agent: Transforming Data into Verifiable Multimodal Storyβ151Jul 5, 2026Updated last month
- β33Jul 3, 2025Updated last year
- Computer-Use Agents as Judges for Generative UIβ45Nov 27, 2025Updated 8 months ago
- β30Apr 16, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams [CVPR'24]β33Jul 23, 2025Updated last year
- [ACL2025 Findings] Benchmarking Multihop Multimodal Internet Agentsβ54Feb 27, 2025Updated last year
- GUICourse: From General Vision Langauge Models to Versatile GUI Agentsβ143Mar 1, 2026Updated 5 months ago
- [NeurIPS 2022] Egocentric Video-Language Pretrainingβ261May 9, 2024Updated 2 years ago
- (ICLR 2025) The Official Code Repository for GUI-World.β69Aug 2, 2026Updated 2 weeks ago
- Official repo of the ICLR 2025 paper "MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos"β28Jul 15, 2025Updated last year
- Transform your CapsLock into an AI key! This AutoHotkey app puts powerful AI capabilities right at your fingertips, supercharging your Wiβ¦β22Oct 31, 2025Updated 9 months ago
- GPT-4V in Wonderland: LMMs as Smartphone Agentsβ134Jul 17, 2024Updated 2 years ago
- Consists of ~500k human annotations on the RICO dataset identifying various icons based on their shapes and semantics, and associations bβ¦β36Jun 27, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [CVPR 2025] GUI-Xplore: Empowering Generalizable GUI Agents with One Explorationβ21Mar 21, 2025Updated last year
- β15Oct 14, 2021Updated 4 years ago
- [ICCV 2023] UniVTG: Towards Unified Video-Language Temporal Groundingβ379May 8, 2024Updated 2 years ago
- ScreenSuite - The most comprehensive benchmarking suite for GUI Agents!β145May 26, 2026Updated 2 months ago
- [ICML'25] Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting | ζ ·ζ¬ηΊ§ε«ηθͺιεΊε€ζ¨‘ειζζΆι΄εΊει’ζ΅β32May 22, 2025Updated last year
- FQGAN: Factorized Visual Tokenization and Generationβ59Mar 29, 2025Updated last year
- β11Sep 22, 2025Updated 10 months ago
- β29Apr 2, 2026Updated 4 months ago
- The model, data and code for the visual GUI Agent SeeClickβ494Jul 13, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Developer project for getting basic API integrations working in under 5 minutesβ11May 22, 2026Updated 2 months ago
- [CVPR 2024] DyBluRF: Dynamic Neural Radiance Fields from Blurry Monocular Videoβ22Jun 14, 2024Updated 2 years ago
- For Ego4D VQ3D Taskβ22Jan 9, 2024Updated 2 years ago
- CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs (CVPR2024)β17Jun 14, 2024Updated 2 years ago
- β16Dec 15, 2025Updated 8 months ago
- Android video semantic segmentation using DeeplabV3+ liteβ10Sep 20, 2019Updated 6 years ago
- AndroidWorld is an environment and benchmark for autonomous agentsβ846Jul 16, 2026Updated 3 weeks ago