Home page for Microsoft Phi-Ground tech-report
☆22Sep 8, 2025Updated last year
Alternatives and similar repositories for Phi-Ground
Users that are interested in Phi-Ground are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆35May 12, 2026Updated 4 months ago
- Official Repo for MageBench: Bridging Large Multimodal Models to Agents☆21Jan 8, 2025Updated last year
- ☆27Mar 26, 2026Updated 6 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- ☆11Sep 20, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- To appear in the 11th International Conference on Learning Representations (ICLR 2023).☆18Feb 24, 2023Updated 3 years ago
- [ACL'25 (Findings)] Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents☆29Feb 17, 2026Updated 7 months ago
- The 1st Spiking Transformer Benchmark (NeurIPS 2025)☆18Jul 10, 2026Updated 3 months ago
- This is a repo to store circuit design datasets☆19Jan 17, 2024Updated 2 years ago
- Pipeline to scrape prompt + image url pairs from LAION `share-dalle-3` discord channel☆11Oct 10, 2023Updated 3 years ago
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms☆25Dec 21, 2025Updated 9 months ago
- Implementation of "PaLM2-VAdapter:" from the multi-modal model paper: "PaLM2-VAdapter: Progressively Aligned Language Model Makes a Stron…☆17Nov 11, 2024Updated last year
- Q-resafe:Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models (ICML'2025)☆16Jun 28, 2025Updated last year
- 2019年全国大学生电子设计大赛G题双路语音调频接收机的FPGA全实现☆17Apr 15, 2020Updated 6 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- implementation of 'The Forward-Forward Algorithm: Some Preliminary Investigations', Hinton 2022☆14Dec 6, 2022Updated 3 years ago
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent w…☆115Sep 8, 2025Updated last year
- syn script for DC Compiler☆15May 15, 2022Updated 4 years ago
- LLMA = LLM + Arithmetic coder, which use LLM to do insane text data compression. LLMA=大模型+算术编码,它能使用LLM对文本数据进行暴力的压缩,达到极高的压缩率。☆22Nov 24, 2024Updated last year
- KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation☆23Apr 23, 2025Updated last year
- Detecting position of sphere of specific radius in point cloud☆11Sep 19, 2018Updated 8 years ago
- ☆15May 13, 2024Updated 2 years ago
- ☆130Oct 3, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- We enable LLM with personalization capability☆11Nov 16, 2023Updated 2 years ago
- Official PyTorch implementation for "Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas", presenting the Merge-Att…☆16Jul 9, 2025Updated last year
- ☆19Mar 8, 2025Updated last year
- understanding of cocotb (In Chinese Only)☆22Jun 10, 2025Updated last year
- A novel approach to the classification of antimicrobial peptides (AMPs) using pre-trained language models to create contextual vectorized…☆17Sep 10, 2024Updated 2 years ago
- ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model☆15Apr 7, 2026Updated 6 months ago
- Here we will track the latest AI Multimodal Models, including Multimodal Foundation Models, LLM, Agent, Audio, Image, Video, Music and 3D…☆36Feb 4, 2025Updated last year
- Official code repository for the paper: "DP-IQA: Utilizing Diffusion Prior for Blind Image Quality Assessment in the Wild"☆25Jun 17, 2025Updated last year
- ☆18Oct 6, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents☆321Aug 24, 2026Updated last month
- Dreamitate: Real-World Visuomotor Policy Learning via Video Generation (CoRL 2024)☆59Jun 7, 2025Updated last year
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost☆122Jul 17, 2025Updated last year
- ☆29Aug 18, 2019Updated 7 years ago
- ☆12Feb 6, 2023Updated 3 years ago
- Dual-Branch Network for Portrait Image Quality Assessment☆19Aug 28, 2026Updated last month
- koa2-lessons☆16Nov 30, 2018Updated 7 years ago