Inference repo for Falcon-Perception and Falcon-OCR model, early-fusion, natively multimodal, dense Autoregressive Transformer models.
☆769Sep 11, 2026Updated this week
Alternatives and similar repositories for Falcon-Perception
Users that are interested in Falcon-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AMoE: Agglomerative Mixture-of-Experts Vision Foundation Models☆59Jun 11, 2026Updated 3 months ago
- Fast state-of-the-art image and video segmentation in portable C/C++☆365Updated this week
- Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized…☆715Apr 14, 2026Updated 5 months ago
- [ECCV2026] X2SAM: Any Segmentation in Images and Videos☆105Jul 13, 2026Updated 2 months ago
- Grounded reasoning agent: Falcon Perception + Gemma 4 VLM on Apple Silicon☆28Apr 5, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning…☆9,469Updated this week
- The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading t…☆11,666Aug 26, 2026Updated 2 weeks ago
- [CVPR 2026 Oral] "INSID3: Training-Free In-Context Segmentation with DINOv3"☆747Jun 26, 2026Updated 2 months ago
- [CVPR2026] Detect Anything via Next Point Prediction☆1,583Feb 22, 2026Updated 6 months ago
- EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept s…☆677Aug 11, 2026Updated last month
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,498Updated this week
- SteerViT is a framework that equips any ViT with the ability to steer both its global and local visual representations with natural langu…☆145Jun 13, 2026Updated 3 months ago
- TIPSv2 (CVPR'26) and TIPS (ICLR'25)☆634Aug 31, 2026Updated 2 weeks ago
- [DEIMv2] Real Time Object Detection Meets DINOv3☆2,027Aug 24, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official code and resources for SAM3-I.☆181Apr 14, 2026Updated 5 months ago
- ☆15Oct 20, 2024Updated last year
- Eagle: Frontier Vision-Language Models with Data-Centric Strategies☆3,562Jun 24, 2026Updated 2 months ago
- Official repository for "AM-RADIO: Reduce All Domains Into One"☆1,948May 29, 2026Updated 3 months ago
- [EMNLP 2026 Main] Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆28Aug 21, 2026Updated 3 weeks ago
- Code for the Molmo2 Vision-Language Model☆727Mar 18, 2026Updated 5 months ago
- ☆228Jun 11, 2026Updated 3 months ago
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction☆3,735Aug 7, 2026Updated last month
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation☆36Aug 13, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- GLM-OCR: Accurate × Fast × Comprehensive☆7,440Apr 21, 2026Updated 4 months ago
- TensorRT Engine for SAM-3 model by Meta AI☆209Jul 2, 2026Updated 2 months ago
- [CVPR 2025] Official PyTorch implementation of "EdgeTAM: On-Device Track Anything Model"☆979Jan 27, 2026Updated 7 months ago
- DINOv3训练示例☆172May 11, 2026Updated 4 months ago
- Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild☆622Jun 1, 2026Updated 3 months ago
- Official code for "No time to train! Training-Free Reference-Based Instance Segmentation"☆317Apr 14, 2026Updated 5 months ago
- Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.☆4,513Aug 22, 2026Updated 3 weeks ago
- Trackers gives you clean, modular re-implementations of leading multi-object tracking algorithms released under the permissive Apache 2.0…☆3,768Updated this week
- MathCode: A Frontier Mathematical Coding Agent☆741Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆96Updated this week
- Reference PyTorch implementation and models for DINOv3☆11,378Jul 15, 2026Updated 2 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,083Aug 18, 2026Updated 3 weeks ago
- [ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs☆101Jan 26, 2026Updated 7 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆851Aug 4, 2026Updated last month
- State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!☆2,363Apr 13, 2026Updated 5 months ago
- [CVPR 2026] InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields☆1,085Apr 3, 2026Updated 5 months ago