Code repo for "Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding"
☆31May 12, 2026Updated 3 months ago
Alternatives and similar repositories for Screen-Point-and-Read
Users that are interested in Screen-Point-and-Read are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Oct 11, 2024Updated last year
- The official repository of "SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World".☆27Jul 27, 2026Updated 3 weeks ago
- [WWW2024 Oral] Harnessing Multi-Role Capabilities of Large Language Models for Open-Domain Question Answering☆15Apr 22, 2025Updated last year
- The dataset includes widget captions that describes UI element's functionalities. It is used for training and evaluation of the widget ca…☆23Jun 24, 2021Updated 5 years ago
- VisionDroid☆22Apr 2, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ORES: Open-vocabulary Responsible Visual Synthesis☆14Dec 12, 2023Updated 2 years ago
- [ECCV 2022] Official pytorch implementation of the paper "FedVLN: Privacy-preserving Federated Vision-and-Language Navigation"☆14Oct 8, 2022Updated 3 years ago
- ☆37May 29, 2025Updated last year
- ☆17Oct 31, 2023Updated 2 years ago
- Official repo for "Imagination-Augmented Natural Language Understanding", NAACL 2022.☆17Aug 30, 2022Updated 3 years ago
- The results and code of our IEEE TCYB 2022 paper, titled "Global-and-Local Collaborative Learning for Co-Salient Object Detection"☆13May 2, 2022Updated 4 years ago
- CVPR25☆28Jul 2, 2025Updated last year
- 吴恩达大模型系列课程中文版,包括《Prompt Engineering》、《Building System》和《LangChain》☆12Jun 7, 2023Updated 3 years ago
- ☆25May 12, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Dataset and models for paper "Game-Based Video-Context Dialogue (EMNLP 2018)"☆19Oct 25, 2018Updated 7 years ago
- Implementation of "Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation"☆27Mar 4, 2021Updated 5 years ago
- Official implementation for "Android in the Zoo: Chain-of-Action-Thought for GUI Agents" (Findings of EMNLP 2024)☆103Oct 14, 2024Updated last year
- ☆17Oct 30, 2023Updated 2 years ago
- Codebase of ACL 2023 Findings "Aerial Vision-and-Dialog Navigation"☆70May 12, 2026Updated 3 months ago
- ☆22Sep 20, 2022Updated 3 years ago
- Implementation of LayoutGAN https://arxiv.org/abs/1901.06767☆17May 12, 2019Updated 7 years ago
- ☆11Sep 18, 2017Updated 8 years ago
- ☆47Mar 19, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Tools for working with the S800 corpus☆12Sep 17, 2020Updated 5 years ago
- ☆15Sep 8, 2025Updated 11 months ago
- The model, data and code for the visual GUI Agent SeeClick☆494Jul 13, 2025Updated last year
- 少年先疯队☆16May 24, 2024Updated 2 years ago
- [EMNLP 2025] Official code for the paper "SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning"☆16May 12, 2026Updated 3 months ago
- Official implementation of "Divide & Bind Your Attention for Improved Generative Semantic Nursing" (BMVC 2023 Oral)☆38Jan 25, 2024Updated 2 years ago
- An image-oriented evaluation tool for image captioning systems (EMNLP-IJCNLP 2019)☆37May 3, 2020Updated 6 years ago
- LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation☆70Aug 9, 2024Updated 2 years ago
- Code for IterInpaint model, presented in Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation (CVPR 2024 work…☆25Jul 21, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of our EMNLP 2022 paper "CPL: Counterfactual Prompt Learning for Vision and Language Models"☆35Dec 5, 2022Updated 3 years ago
- ☆21Apr 2, 2024Updated 2 years ago
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models (ICLR2026)☆23Jun 24, 2026Updated last month
- Cornell Instruction Following Framework☆35Oct 11, 2021Updated 4 years ago
- A PyTorch implementation of SSCR☆23Aug 12, 2024Updated 2 years ago
- ☆19Jul 18, 2024Updated 2 years ago
- Official code for paper "Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning"☆15Jun 12, 2025Updated last year