Home page for Microsoft Phi-Ground tech-report
☆22Sep 8, 2025Updated last year
Alternatives and similar repositories for Phi-Ground
Users that are interested in Phi-Ground are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repo for MageBench: Bridging Large Multimodal Models to Agents☆21Jan 8, 2025Updated last year
- This repo is for CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering.☆14Mar 6, 2024Updated 2 years ago
- A simple implementation of reverse mode automatic differentiation in C++ without the use of any libraries.☆13Jul 3, 2018Updated 8 years ago
- LMM for VQA, tcsvt version☆10Jul 19, 2024Updated 2 years ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆11Jan 17, 2021Updated 5 years ago
- ☆16Aug 5, 2025Updated last year
- Quickly hashing all subexpressions of a program modulo alpha-renaming☆17Sep 7, 2021Updated 5 years ago
- ☆10Dec 9, 2021Updated 4 years ago
- [ACL'25 (Findings)] Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents☆29Feb 17, 2026Updated 7 months ago
- [ECML 2020] OBProx-SG☆16Mar 23, 2021Updated 5 years ago
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms☆25Dec 21, 2025Updated 8 months ago
- Implementation of Balanced Graph Partitioning Konstantin" - Andreev and Harald Racke (Authors of the paper) by Ivan Vigorito and Lorenzo …☆14Feb 17, 2023Updated 3 years ago
- Implementation of "PaLM2-VAdapter:" from the multi-modal model paper: "PaLM2-VAdapter: Progressively Aligned Language Model Makes a Stron…☆17Nov 11, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [NeurIPS'25] GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents☆416Apr 13, 2026Updated 5 months ago
- [ICLR 2026] Efficient Agent Training for Computer Use☆145Sep 5, 2025Updated last year
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent w…☆115Sep 8, 2025Updated last year
- KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation☆23Apr 23, 2025Updated last year
- [ICLR ML4RS 2025] Official implementation for the paper "Tackling Few-Shot Segmentation in Remote Sensing via Inpainting Diffusion Model"☆17Feb 2, 2026Updated 7 months ago
- ☆129Oct 3, 2025Updated 11 months ago
- We enable LLM with personalization capability☆11Nov 16, 2023Updated 2 years ago
- This is the official code for the paper "Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning" (NeurIPS2024)☆28Sep 10, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Training codebase for K2-V2☆22Dec 17, 2025Updated 9 months ago
- [ICCV 2025] MMReason, MLLMs, step by step, reasoning benchmark, AGI☆15Apr 25, 2026Updated 4 months ago
- Official PyTorch implementation for "Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas", presenting the Merge-Att…☆16Jul 9, 2025Updated last year
- Here we will track the latest AI Multimodal Models, including Multimodal Foundation Models, LLM, Agent, Audio, Image, Video, Music and 3D…☆36Feb 4, 2025Updated last year
- Code release for "Memorization in 3D Shape Generation: An Empirical Study"☆21Dec 30, 2025Updated 8 months ago
- Official code repository for the paper: "DP-IQA: Utilizing Diffusion Prior for Blind Image Quality Assessment in the Wild"☆25Jun 17, 2025Updated last year
- Analysis of video quality datasets via design of minimalistic video quality models☆24Jul 15, 2024Updated 2 years ago
- ☆18Oct 6, 2022Updated 3 years ago
- [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents☆318Aug 24, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Dreamitate: Real-World Visuomotor Policy Learning via Video Generation (CoRL 2024)☆59Jun 7, 2025Updated last year
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost☆122Jul 17, 2025Updated last year
- A PDDL Solver in C++.☆15Jan 5, 2024Updated 2 years ago
- ☆12Feb 6, 2023Updated 3 years ago
- 邮件发送平台,生产平台☆17Jul 17, 2019Updated 7 years ago
- Crawl data from articles of the New York Times website☆13Oct 23, 2019Updated 6 years ago
- [NeurIPS 2025] HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation☆78Sep 19, 2025Updated last year