Inference repo for Falcon-Perception and Falcon-OCR model, early-fusion, natively multimodal, dense Autoregressive Transformer models.
☆778Sep 11, 2026Updated 3 weeks ago
Alternatives and similar repositories for Falcon-Perception
Users that are interested in Falcon-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AMoE: Agglomerative Mixture-of-Experts Vision Foundation Models☆60Jun 11, 2026Updated 3 months ago
- Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized…☆725Apr 14, 2026Updated 5 months ago
- [ECCV2026] X2SAM: Any Segmentation in Images and Videos☆106Jul 13, 2026Updated 2 months ago
- Grounded reasoning agent: Falcon Perception + Gemma 4 VLM on Apple Silicon☆28Apr 5, 2026Updated 6 months ago
- RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning…☆9,697Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading t…☆11,867Sep 18, 2026Updated 2 weeks ago
- [CVPR 2026 Oral] "INSID3: Training-Free In-Context Segmentation with DINOv3"☆766Jun 26, 2026Updated 3 months ago
- [CVPR2026] Detect Anything via Next Point Prediction☆1,600Feb 22, 2026Updated 7 months ago
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,565Updated this week
- TIPSv2 (CVPR'26) and TIPS (ICLR'25)☆638Aug 31, 2026Updated last month
- [DEIMv2] Real Time Object Detection Meets DINOv3☆2,049Aug 24, 2026Updated last month
- Official code and resources for SAM3-I.☆183Apr 14, 2026Updated 5 months ago
- ☆15Oct 20, 2024Updated last year
- Eagle: Frontier Vision-Language Models with Data-Centric Strategies☆3,679Jun 24, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official repository for "AM-RADIO: Reduce All Domains Into One"☆1,967Updated this week
- [EMNLP 2026 Main] Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆28Aug 21, 2026Updated last month
- Few Shot Semantic Segmentation Meets SAM3☆52May 21, 2026Updated 4 months ago
- Code for the Molmo2 Vision-Language Model☆737Mar 18, 2026Updated 6 months ago
- ☆95Apr 7, 2026Updated 5 months ago
- ☆235Jun 11, 2026Updated 3 months ago
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction☆3,784Aug 7, 2026Updated last month
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation☆36Aug 13, 2026Updated last month
- GLM-OCR: Accurate × Fast × Comprehensive☆7,485Apr 21, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- TensorRT Engine for SAM-3 model by Meta AI☆219Jul 2, 2026Updated 3 months ago
- [CVPR 2025] Official PyTorch implementation of "EdgeTAM: On-Device Track Anything Model"☆990Jan 27, 2026Updated 8 months ago
- DINOv3训练示例☆173Updated this week
- Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild☆625Jun 1, 2026Updated 4 months ago
- Official code for "No time to train! Training-Free Reference-Based Instance Segmentation"☆318Apr 14, 2026Updated 5 months ago
- Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.☆4,526Sep 29, 2026Updated last week
- Reference PyTorch implementation and models for DINOv3☆11,486Jul 15, 2026Updated 2 months ago
- Trackers gives you clean, modular re-implementations of leading multi-object tracking algorithms released under the permissive Apache 2.0…☆3,905Updated this week
- MathCode: A Frontier Mathematical Coding Agent☆744Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆102Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,137Aug 18, 2026Updated last month
- [ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs☆102Jan 26, 2026Updated 8 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆854Aug 4, 2026Updated 2 months ago
- State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!☆2,381Apr 13, 2026Updated 5 months ago
- YOLOE: Real-Time Seeing Anything [ICCV 2025]☆2,302Jun 26, 2025Updated last year
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 5 months ago