Clone of DeepSeek Thinking-with-Visual-Primitives
☆181Apr 30, 2026Updated 3 months ago
Alternatives and similar repositories for Thinking-with-Visual-Primitives
Users that are interested in Thinking-with-Visual-Primitives are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Mar 24, 2026Updated 4 months ago
- Holistic Evaluation of Multimodal LLMs on Spatial Intelligence☆120Jul 1, 2026Updated last month
- Spatial Aptitude Training for Multimodal Langauge Models☆33Feb 8, 2026Updated 6 months ago
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery☆15Feb 1, 2026Updated 6 months ago
- 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning☆15Jun 2, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Functional AWK. Experiment, no quality guarantees. Not aiming to be compatible with the AWK standard.☆29Dec 9, 2025Updated 8 months ago
- [ICCV 2025] D^3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection☆17Jul 11, 2026Updated last month
- [ECCV 2026 Oral] Official implementation of "Make Geometry Matter for Spatial Reasoning"☆54Aug 6, 2026Updated last week
- ☆10Apr 7, 2025Updated last year
- Public repository for the Remote Labor Index (RLI)☆75Nov 3, 2025Updated 9 months ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, ex…☆25Aug 3, 2026Updated last week
- ☆15Nov 25, 2022Updated 3 years ago
- [ACM MM'26] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation☆64May 14, 2026Updated 3 months ago
- [ACM MM-24] Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization☆13Oct 8, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ECCV 26'] Official codebase for the paper LaViT☆35Jul 30, 2026Updated 2 weeks ago
- [NeurIPS 2025] Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing☆97Jul 27, 2025Updated last year
- [CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"☆88May 12, 2026Updated 3 months ago
- ☆22Dec 18, 2025Updated 7 months ago
- Probabilistic Jacobian-based Saliency Maps Attacks☆19Nov 28, 2020Updated 5 years ago
- Learning Debiased and Disentangled Representations for Semantic Segmentation (NeurIPS 2021)☆13Jan 23, 2022Updated 4 years ago
- [CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"☆216Mar 19, 2026Updated 4 months ago
- ☆17Nov 20, 2024Updated last year
- Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆27Apr 5, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆126Apr 9, 2026Updated 4 months ago
- [NeurIPS 2025] The official PyTorch implementation of the "Vision Function Layer in MLLM".☆33Dec 18, 2025Updated 7 months ago
- Official Github of "Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework"☆22Jan 4, 2026Updated 7 months ago
- This repository is the official implementation of "Look-Back: Implicit Visual Re-focusing in MLLM Reasoning".☆99Jul 10, 2025Updated last year
- QuoteSum is a textual QA dataset containing Semi-Extractive Multi-source Question Answering (SEMQA) examples written by humans, based on …☆13Mar 25, 2024Updated 2 years ago
- [CVPR'25] CoMatcher: Multi-View Collaborative Feature Matching☆23Aug 21, 2025Updated 11 months ago
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning☆19Jun 2, 2026Updated 2 months ago
- [World-Model-Survey-2024] Paper list and projects for World Model☆15Oct 31, 2024Updated last year
- Minimum viable code for the Decodable Information Bottleneck paper. Pytorch Implementation.☆12Oct 20, 2020Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Cross-Modal RGB-Event Benchmark for Multi-Object Tracking and Detection.☆13Oct 17, 2023Updated 2 years ago
- Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"☆40Updated this week
- Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.1983…☆100Mar 9, 2026Updated 5 months ago
- CTF Challenge for CSAW Finals 2021☆13Nov 17, 2021Updated 4 years ago
- Official Pytorch Implementation of Paper "DarwinLM: Evolutionary Structured Pruning of Large Language Models"☆20Feb 21, 2025Updated last year
- Official code for paper "Reasoning Fails Where Step Flow Breaks" (ACL 2026)☆19Apr 19, 2026Updated 3 months ago
- Pytorch Implementation of Reliable Thinking with Images.☆28May 3, 2026Updated 3 months ago