Inference repo for Falcon-Perception and Falcon-OCR model, early-fusion, natively multimodal, dense Autoregressive Transformer models.
☆756Aug 13, 2026Updated last week
Alternatives and similar repositories for Falcon-Perception
Users that are interested in Falcon-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AMoE: Agglomerative Mixture-of-Experts Vision Foundation Models☆55Jun 11, 2026Updated 2 months ago
- Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized…☆704Apr 14, 2026Updated 4 months ago
- [ECCV2026] X2SAM: Any Segmentation in Images and Videos☆100Jul 13, 2026Updated last month
- Grounded reasoning agent: Falcon Perception + Gemma 4 VLM on Apple Silicon☆28Apr 5, 2026Updated 4 months ago
- RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning…☆9,053Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading t…☆11,469Aug 14, 2026Updated last week
- [CVPR 2026 Oral] "INSID3: Training-Free In-Context Segmentation with DINOv3"☆732Jun 26, 2026Updated 2 months ago
- [CVPR2026] Detect Anything via Next Point Prediction☆1,560Feb 22, 2026Updated 6 months ago
- SteerViT is a framework that equips any ViT with the ability to steer both its global and local visual representations with natural langu…☆122Jun 13, 2026Updated 2 months ago
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,412Updated this week
- TIPSv2 (CVPR'26) and TIPS (ICLR'25)☆603Jun 1, 2026Updated 2 months ago
- [DEIMv2] Real Time Object Detection Meets DINOv3☆1,997Updated this week
- Official code and resources for SAM3-I.☆178Apr 14, 2026Updated 4 months ago
- Official repository for "AM-RADIO: Reduce All Domains Into One"☆1,932May 29, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [EMNLP 2026 Main] Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆27Updated this week
- Few Shot Semantic Segmentation Meets SAM3☆51May 21, 2026Updated 3 months ago
- Code for the Molmo2 Vision-Language Model☆714Mar 18, 2026Updated 5 months ago
- ☆94Apr 7, 2026Updated 4 months ago
- ☆213Jun 11, 2026Updated 2 months ago
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction☆3,675Aug 7, 2026Updated 2 weeks ago
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation☆26Aug 13, 2026Updated last week
- TensorRT Engine for SAM-3 model by Meta AI☆200Jul 2, 2026Updated last month
- GLM-OCR: Accurate × Fast × Comprehensive☆7,360Apr 21, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025] Official PyTorch implementation of "EdgeTAM: On-Device Track Anything Model"☆971Jan 27, 2026Updated 6 months ago
- DINOv3训练示例☆171May 11, 2026Updated 3 months ago
- Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild☆608Jun 1, 2026Updated 2 months ago
- Official code for "No time to train! Training-Free Reference-Based Instance Segmentation"☆316Apr 14, 2026Updated 4 months ago
- Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.☆4,431Updated this week
- Trackers gives you clean, modular re-implementations of leading multi-object tracking algorithms released under the permissive Apache 2.0…☆3,713Updated this week
- MathCode: A Frontier Mathematical Coding Agent☆732Updated this week
- ☆87Updated this week
- Reference PyTorch implementation and models for DINOv3☆11,243Jul 15, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,941Aug 18, 2026Updated last week
- [ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs☆100Jan 26, 2026Updated 7 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆845Aug 4, 2026Updated 3 weeks ago
- State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!☆2,353Apr 13, 2026Updated 4 months ago
- [CVPR 2026] InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields☆1,073Apr 3, 2026Updated 4 months ago
- Vero: An Open RL Recipe for General Visual Reasoning☆144Jun 19, 2026Updated 2 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 4 months ago