☆21Jan 22, 2026Updated 7 months ago
Alternatives and similar repositories for UITron-Speech
Users that are interested in UITron-Speech are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆66Sep 6, 2025Updated 11 months ago
- [CVPR 2026] Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR☆19Mar 23, 2026Updated 5 months ago
- OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models☆30Feb 4, 2026Updated 6 months ago
- Repository for the paper "InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners"☆67Dec 4, 2025Updated 8 months ago
- ☆34Sep 19, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16Aug 27, 2025Updated last year
- [ACL 2026] Closing the Modality Reasoning Gap for Speech Large Language Models☆18Apr 17, 2026Updated 4 months ago
- [ACL 2025] Research code for the paper "OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents"☆23Jun 19, 2025Updated last year
- [ICCV 2025 Highlight] Less is More: Empowering GUI Agent with Context-Aware Simplification☆48Mar 12, 2026Updated 5 months ago
- Source code of the paper "V-Droid: Advancing Mobile GUI Agent Through Generative Verifiers"☆35Feb 2, 2026Updated 6 months ago
- ☆18Oct 27, 2025Updated 10 months ago
- [NeurIPS 2025]"Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning"☆108Oct 21, 2025Updated 10 months ago
- finetune your florence2 model easy☆21Jul 27, 2024Updated 2 years ago
- This is a Python implementation of detecting mobile phones while driving using YOLO v5, performed using Kaggle's State Farm Distracted Dr…☆10Dec 28, 2021Updated 4 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ACL 2025] GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent☆68May 28, 2025Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost☆121Jul 17, 2025Updated last year
- Native End-to-End Full-Duplex Spoken Language Model☆128Aug 18, 2026Updated last week
- [NeurIPS 2025] UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents☆61Nov 27, 2025Updated 9 months ago
- ☆10Apr 22, 2021Updated 5 years ago
- ☆17May 14, 2025Updated last year
- ☆29Jul 4, 2025Updated last year
- A simple visual test-time scaling method for GUI agent grounding☆26Dec 7, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆81Sep 3, 2025Updated 11 months ago
- [ICLR 2020] Haotao Wang, Tianlong Chen, Zhangyang Wang, Kede Ma, "I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifie…☆20Dec 30, 2021Updated 4 years ago
- Scene Parsing via Integrated Classification Model and Variance-Based Regularization (Matlab&Caffe), In CVPR 2019☆12Jun 11, 2019Updated 7 years ago
- Official repository for paper Auto-scaling Continuous Memory for GUI Agent☆30Feb 2, 2026Updated 6 months ago
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- [MM'23] ProTegO: Protect Text Content against OCR Extraction Attack☆14Mar 12, 2024Updated 2 years ago
- OS-ATLAS: A Foundation Action Model For Generalist GUI Agents☆453Apr 20, 2025Updated last year
- 2022 WAIC 黑客松蚂蚁财富赛道:AntSQL大规模金融语义解析中文Text-to-SQL挑战赛 一位萌新的代码 嘻嘻嘻☆14Mar 11, 2023Updated 3 years ago
- Official implementation of VLAA-GUI series☆36Jun 20, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code for Semantic Adversarial Attacks☆11Oct 12, 2021Updated 4 years ago
- A Pytorch version of LPCNet, including dump weight☆36May 5, 2022Updated 4 years ago
- OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video☆20Apr 24, 2026Updated 4 months ago
- ☆42Aug 23, 2026Updated last week
- [Pattern Recognition'24] Looking Beyond Input Frames: Self-Supervised Adaptation for Video Super-Resolution☆16Apr 1, 2024Updated 2 years ago
- Database of "Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model", ECCV 2020☆13May 2, 2022Updated 4 years ago
- [ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.☆30Jul 2, 2026Updated last month