[NeurIPS2025 Spotlight ๐ฅ ] Official implementation of ๐ธ "UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface"
โ284Nov 5, 2025Updated 10 months ago
Alternatives and similar repositories for UFO
Users that are interested in UFO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"โ639Jan 17, 2026Updated 7 months ago
- code for the paper "CoReS: Orchestrating the Dance of Reasoning and Segmentation"โ23Nov 24, 2025Updated 9 months ago
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Modelโ97Jul 17, 2025Updated last year
- [ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learningโ352Feb 9, 2026Updated 6 months ago
- Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)โ1,666Aug 4, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer โข AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR2025] Project for "HyperSeg: Towards Universal Visual Segmentation with Large Language Model".โ184Dec 13, 2024Updated last year
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generationโ29May 27, 2025Updated last year
- [ECCV2024 Oral๐ฅ] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"โ364Jan 14, 2025Updated last year
- [ICLR2025] Text4Seg: Reimagining Image Segmentation as Text Generationโ176Nov 8, 2025Updated 10 months ago
- MLLMSeg: Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoderโ57Jun 12, 2026Updated 2 months ago
- โ12Nov 26, 2024Updated last year
- ๐ฎ UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning (NeurIPS 2025)โ252Jan 4, 2026Updated 8 months ago
- Boosting 3D Object Detection via Object-Focused Image Fusionโ59Sep 11, 2022Updated 3 years ago
- A curated list of publications on image and video segmentation leveraging Multimodal Large Language Models (MLLMs), highlighting state-ofโฆโ234Aug 29, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient โข AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code release for "SegLLM: Multi-round Reasoning Segmentation"โ129Feb 20, 2025Updated last year
- [ECCV2024] This is an official implementation for "PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model"โ272Dec 30, 2024Updated last year
- Official code for DeepSound-V1โ12May 14, 2025Updated last year
- [CVPR 2025] Official Pytorch Code for Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentationโ50Mar 27, 2025Updated last year
- [ICLR 2025] Official Pytorch Implementation of MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmโฆโ28Apr 3, 2025Updated last year
- [ICCV 2025] Official implementation of "InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models"โ56Feb 10, 2025Updated last year
- [NeurIPS 2025] Official code implementation of Perception R1: Pioneering Perception Policy with Reinforcement Learningโ290Jul 15, 2025Updated last year
- โ45Jul 9, 2025Updated last year
- [ICLR 2026 ๐ฅ ] Official implementation of "UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing"โ152Jan 26, 2026Updated 7 months ago
- Simple, predictable pricing with DigitalOcean hosting โข AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'โโ2,276Oct 29, 2025Updated 10 months ago
- [ACL'25 Main] Official Implementation of HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Languagโฆโ56Jun 1, 2026Updated 3 months ago
- [CVPR2026] Detect Anything via Next Point Predictionโ1,574Feb 22, 2026Updated 6 months ago
- Rui Qian, Xin Yin, Dejing Douโ : Reasoning to Attend: Try to Understand How <SEG> Token Works (CVPR 2025)โ55Feb 4, 2026Updated 7 months ago
- โ22Jan 9, 2026Updated 7 months ago
- โ22Jun 15, 2023Updated 3 years ago
- [CVPR 2024] PixelLM is an effective and efficient LMM for pixel-level reasoning and understanding.โ274Feb 11, 2025Updated last year
- [ECCV 2024] The official code of paper "Open-Vocabulary SAM".โ1,033Aug 4, 2025Updated last year
- Code for ChatRex: Taming Multimodal LLM for Joint Perception and Understandingโ216Oct 15, 2025Updated 10 months ago
- Open source password manager - Proton Pass โข AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ECCV 2024] SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentationโ52Mar 20, 2025Updated last year
- Project Page for "LISA: Reasoning Segmentation via Large Language Model"โ2,674Feb 16, 2025Updated last year
- [ICCV 2025] GroundingSuite: Measuring Complex Multi-Granular Pixel Groundingโ77Jun 26, 2025Updated last year
- Solve Visual Understanding with Reinforced VLMsโ6,019Jul 7, 2026Updated 2 months ago
- [CVPR 2025] Official PyTorch Implementation of GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentaโฆโ70Jun 23, 2025Updated last year
- [CVPR 2024 ๐ฅ] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses thaโฆโ967Updated this week
- [ICCV 2025] MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentationโ23Sep 5, 2025Updated last year