Inference repo for Falcon-Perception and Falcon-OCR model, early-fusion, natively multimodal, dense Autoregressive Transformer models.
☆742Jul 14, 2026Updated 3 weeks ago
Alternatives and similar repositories for Falcon-Perception
Users that are interested in Falcon-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AMoE: Agglomerative Mixture-of-Experts Vision Foundation Models☆53Jun 11, 2026Updated last month
- Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized…☆693Apr 14, 2026Updated 3 months ago
- Grounded reasoning agent: Falcon Perception + Gemma 4 VLM on Apple Silicon☆28Apr 5, 2026Updated 3 months ago
- RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning…☆8,858Updated this week
- The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading t…☆11,188Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR 2026 Oral] "INSID3: Training-Free In-Context Segmentation with DINOv3"☆706Jun 26, 2026Updated last month
- [CVPR2026] Detect Anything via Next Point Prediction☆1,531Feb 22, 2026Updated 5 months ago
- SteerViT is a framework that equips any ViT with the ability to steer both its global and local visual representations with natural langu…☆118Jun 13, 2026Updated last month
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,291Updated this week
- TIPSv2 (CVPR'26) and TIPS (ICLR'25)☆576Jun 1, 2026Updated 2 months ago
- [DEIMv2] Real Time Object Detection Meets DINOv3☆1,967Mar 24, 2026Updated 4 months ago
- Official code and resources for SAM3-I.☆175Apr 14, 2026Updated 3 months ago
- ☆15Oct 20, 2024Updated last year
- Official repository for "AM-RADIO: Reduce All Domains Into One"☆1,909May 29, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆27Apr 5, 2026Updated 3 months ago
- Code for the Molmo2 Vision-Language Model☆701Mar 18, 2026Updated 4 months ago
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction☆3,616Jul 17, 2026Updated 2 weeks ago
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation☆26Jun 29, 2026Updated last month
- TensorRT Engine for SAM-3 model by Meta AI☆193Jul 2, 2026Updated last month
- GLM-OCR: Accurate × Fast × Comprehensive☆7,245Apr 21, 2026Updated 3 months ago
- [CVPR 2025] Official PyTorch implementation of "EdgeTAM: On-Device Track Anything Model"☆955Jan 27, 2026Updated 6 months ago
- DINOv3训练示例☆169May 11, 2026Updated 2 months ago
- Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild☆604Jun 1, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official code for "No time to train! Training-Free Reference-Based Instance Segmentation"☆315Apr 14, 2026Updated 3 months ago
- Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.☆4,402Jul 23, 2026Updated last week
- Trackers gives you clean, modular re-implementations of leading multi-object tracking algorithms released under the permissive Apache 2.0…☆3,568Updated this week
- MathCode: A Frontier Mathematical Coding Agent☆585Jun 15, 2026Updated last month
- ☆82Updated this week
- Reference PyTorch implementation and models for DINOv3☆11,093Jul 15, 2026Updated 2 weeks ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,562May 10, 2026Updated 2 months ago
- [ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs☆100Jan 26, 2026Updated 6 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆836Jul 14, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!☆2,332Apr 13, 2026Updated 3 months ago
- [CVPR 2026] InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields☆1,059Apr 3, 2026Updated 4 months ago
- Vero: An Open RL Recipe for General Visual Reasoning☆140Jun 19, 2026Updated last month
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 3 months ago
- YOLOE: Real-Time Seeing Anything [ICCV 2025]☆2,224Jun 26, 2025Updated last year
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs☆329Jun 18, 2026Updated last month
- [AAAI2026] X-SAM: From Segment Anything to Any Segmentation☆386Jul 14, 2026Updated 3 weeks ago