Official code for Attention-driven GUI Grounding, AAAI2025
☆16Dec 17, 2024Updated last year
Alternatives and similar repositories for TAG
Users that are interested in TAG are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the multi-agent computer use project.☆21Jul 3, 2026Updated last month
- ☆35Jun 20, 2024Updated 2 years ago
- ☆12Dec 20, 2024Updated last year
- ☆11Oct 2, 2024Updated last year
- The code implementation of GraCeFul (Accepted in COLING 2025)☆13Jan 27, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- ☆13Jul 22, 2026Updated last month
- On the Robustness of GUI Grounding Models Against Image Attacks☆12Apr 8, 2025Updated last year
- Code for "ATTA: Anomaly-aware Test-Time Adaptation for Out-of-Distribution Detection in Segmentation" (NeurIPS 23)☆16Apr 12, 2024Updated 2 years ago
- ☆11Aug 20, 2025Updated last year
- Memory footprint reduction for transformer models☆11Jan 24, 2023Updated 3 years ago
- GUICourse: From General Vision Langauge Models to Versatile GUI Agents☆142Mar 1, 2026Updated 5 months ago
- Pre-trained grapheme-to-phoneme (G2P) models☆26Jul 27, 2021Updated 5 years ago
- Implementation for paper "Link Prediction on Heterophilic Graphs via Disentangled Representation Learning"☆13Aug 26, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of the paper MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction☆20Feb 19, 2026Updated 6 months ago
- PoC for Paper: BunnyHop Exploiting the Instruction Prefetcher (USENIX Security 2023)☆14Aug 17, 2023Updated 3 years ago
- something for paper agent☆11Dec 18, 2024Updated last year
- [CVPR 2024] Code and datasets for 'Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos'☆14Jun 16, 2024Updated 2 years ago
- [NeurIPS'25] GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents☆415Apr 13, 2026Updated 4 months ago
- ☆10Mar 11, 2022Updated 4 years ago
- Controllable mage captioning model with unsupervised modes☆21Apr 14, 2023Updated 3 years ago
- [AAAI'26] Official implementation of CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augm…☆11Dec 5, 2025Updated 8 months ago
- Making an Ace Combat-Style flight shooting game (only one mission) with Unity.☆16Jan 31, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The code of "Image-text Retrieval via Preserving Main Semantic of Vision" in ICME 2023.☆15Dec 25, 2023Updated 2 years ago
- This is the repository for paper EscapeBench: Pushing Language Models to Think Outside the Box☆18Dec 19, 2024Updated last year
- [EMNLP 2024] SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information☆11Oct 11, 2024Updated last year
- Reproduction code for paper "MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft"☆21Jun 12, 2026Updated 2 months ago
- ☆18May 25, 2023Updated 3 years ago
- SPA: Efficient User-Preference Alignment against Uncertainty in Medical Image Segmentation (ICCV 2025)☆16Sep 26, 2025Updated 11 months ago
- [NAACL 2025] Guiding Large Language Models in Code Execution with Fine-grained Multimodal Chain-of-Thought Reasoning☆10Feb 9, 2025Updated last year
- Benchmark for Anomaly Detection in Semantic Segmentation☆12Feb 27, 2026Updated 6 months ago
- CheXficient☆15Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- simple solution based on Gradient Boost and Random Forest, rank 24/3251 (top 1%) within 60 lines of python code☆14Jun 21, 2019Updated 7 years ago
- ☆16Mar 12, 2025Updated last year
- [CVPR'26, Findings] AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting☆15May 18, 2026Updated 3 months ago
- Membership Inference Attack against Graph Neural Networks☆12Nov 9, 2022Updated 3 years ago
- [AAAI 2024] MESED: A Multi-modal Entity Set Expansion Dataset with Fine-grained Semantic Classes and Hard Negative Entities☆15Apr 26, 2024Updated 2 years ago
- The official repo for "Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation", ECCV 2024☆18Oct 11, 2024Updated last year
- ☆14May 23, 2022Updated 4 years ago