☆21Jan 22, 2026Updated 8 months ago
Alternatives and similar repositories for UITron-Speech
Users that are interested in UITron-Speech are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner☆24Aug 7, 2025Updated last year
- [CVPR 2026] Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR☆19Mar 23, 2026Updated 6 months ago
- OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models☆29Feb 4, 2026Updated 8 months ago
- Repository for the paper "InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners"☆66Dec 4, 2025Updated 10 months ago
- ☆34Sep 19, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ACL 2026] Closing the Modality Reasoning Gap for Speech Large Language Models☆18Apr 17, 2026Updated 5 months ago
- ☆19Aug 27, 2025Updated last year
- [ACL 2025] Research code for the paper "OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents"☆24Jun 19, 2025Updated last year
- [ICCV 2025 Highlight] Less is More: Empowering GUI Agent with Context-Aware Simplification☆48Mar 12, 2026Updated 6 months ago
- Source code of the paper "V-Droid: Advancing Mobile GUI Agent Through Generative Verifiers"☆41Feb 2, 2026Updated 8 months ago
- ☆18Oct 27, 2025Updated 11 months ago
- [NeurIPS 2025]"Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning"☆110Oct 21, 2025Updated 11 months ago
- finetune your florence2 model easy☆21Jul 27, 2024Updated 2 years ago
- LAVIS - A One-stop Library for Language-Vision Intelligence☆48Aug 5, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ACL 2025] GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent☆68May 28, 2025Updated last year
- (CVPR 2026) Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion☆17Mar 8, 2026Updated 7 months ago
- Graph Convolutional Module for Temporal Action Localization in Videos☆10Jul 4, 2020Updated 6 years ago
- A list of widely-used open-sourced autoregressive or non-autoregressive TTS models☆21Apr 13, 2026Updated 5 months ago
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost☆122Jul 17, 2025Updated last year
- Native End-to-End Full-Duplex Spoken Language Model☆152Aug 18, 2026Updated last month
- ☆10Apr 22, 2021Updated 5 years ago
- ☆18May 14, 2025Updated last year
- Offical code repository of ”DAAD: Dynamic Analysis and Adaptive Discriminator for Fake News Detection“☆22Aug 22, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆29Jul 4, 2025Updated last year
- ☆10Mar 22, 2022Updated 4 years ago
- [CVPR 2026] GThinker, Reasoning MLLM, Visual Cues, Visual Rethinking☆18Mar 9, 2026Updated 6 months ago
- ☆80Sep 3, 2025Updated last year
- ScreenQA dataset was introduced in the "ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots" paper. It contains ~86K …☆152Feb 7, 2025Updated last year
- A real-time inferencing of multistreaming YOWOv3(Spatio Temporal Action Detection task) using (UCF101-24) dataset. The repo is extension …☆27May 15, 2026Updated 4 months ago
- Scene Parsing via Integrated Classification Model and Variance-Based Regularization (Matlab&Caffe), In CVPR 2019☆12Jun 11, 2019Updated 7 years ago
- Official repository for paper Auto-scaling Continuous Memory for GUI Agent☆30Feb 2, 2026Updated 8 months ago
- ☆14Dec 12, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [CVPR 2024] Face2Diffusion for Fast and Editable Face Personalization https://arxiv.org/abs/2403.05094☆96Mar 28, 2024Updated 2 years ago
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- rkllm_talking is a standalone compiled voice communication system based on a large model || rkllm_talking 是一个独立编译的基于大模…☆15Oct 13, 2024Updated last year
- [MM'23] ProTegO: Protect Text Content against OCR Extraction Attack☆14Mar 12, 2024Updated 2 years ago
- OS-ATLAS: A Foundation Action Model For Generalist GUI Agents☆456Apr 20, 2025Updated last year
- Codes for our ICLR2020 paper: Knowledge Consistency between Neural Networks and Beyond☆16Jan 11, 2020Updated 6 years ago
- PersoSim - the open source eID simulator☆16May 21, 2026Updated 4 months ago