official repo for `thinking with images through-self-calling`
☆26Dec 28, 2025Updated 7 months ago
Alternatives and similar repositories for think-with-images-through-self-calling
Users that are interested in think-with-images-through-self-calling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repo of From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs☆24Jun 23, 2026Updated last month
- [ICLR 2026] Geometric-Mean Policy Optimization☆104Jan 26, 2026Updated 6 months ago
- [NeurIPS 2024] Artemis: Towards Referential Understanding in Complex Videos☆27Apr 8, 2025Updated last year
- [ICML 2026 Spotlight] Code for miXed Discrete Diffusion Language Model☆29Mar 16, 2026Updated 5 months ago
- ☆31Sep 24, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Implementation of paper "CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis"☆28Dec 19, 2025Updated 7 months ago
- [CVPR 2026] LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding☆52Jul 7, 2026Updated last month
- [CVPR 2025] DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution☆59Mar 4, 2025Updated last year
- ☆20May 15, 2026Updated 3 months ago
- [CVPR 2026] Thinking with Programming Vision: Towards a Unified View for Thinking with Images☆72Jan 23, 2026Updated 6 months ago
- ☆23Jul 4, 2026Updated last month
- ☆1,265Nov 20, 2025Updated 8 months ago
- ☆35Oct 8, 2025Updated 10 months ago
- vHeat: Building Vision Models upon Heat Conduction☆285Jun 12, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official PyTorch Implementation☆17Dec 3, 2022Updated 3 years ago
- [ICLR 2026] Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization☆32Mar 6, 2026Updated 5 months ago
- [CVPR 2025] Adaptive Keyframe Sampling for Long Video Understanding☆231Dec 19, 2025Updated 7 months ago
- ☆26Feb 27, 2022Updated 4 years ago
- Official repo for [ICML 2026] "Position: The Systemic Lack of Agency in Visual Reasoning"☆31Jun 17, 2026Updated 2 months ago
- Awesome OVD-OVS - A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future☆220Apr 3, 2025Updated last year
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 8 months ago
- (CVPR2023/TPAMI2024) Integrally Pre-Trained Transformer Pyramid Networks -- A Hierarchical Vision Transformer for Masked Image Modeling☆216Jul 28, 2024Updated 2 years ago
- [ICCV 2023] Generative Prompt Model for Weakly Supervised Object Localization☆57Nov 10, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [AAAI2025] ChatterBox: Multi-round Multimodal Referring and Grounding, Multimodal, Multi-round dialogues☆62May 2, 2025Updated last year
- Fully open reproduction of DeepSeek-R1☆11Mar 24, 2025Updated last year
- ☆10Feb 21, 2023Updated 3 years ago
- Deepest Season 6 Meta-Learning study papers plus alpha☆25Mar 4, 2020Updated 6 years ago
- RegionReasoner: Region-Grounded Multi-Round Visual Reasoning (ICLR 2026)☆21Jun 12, 2026Updated 2 months ago
- ☆16Updated this week
- [CVPR 2026 Main] MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation☆29Aug 4, 2026Updated 2 weeks ago
- Representation Surgery for Multi-Task Model Merging. ICML, 2024.☆49Oct 10, 2024Updated last year
- ☆13May 17, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- official impelmentation of Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input☆67Aug 30, 2024Updated last year
- GOlang server-side library enabling the creation of custom management dashboards☆16Jan 31, 2017Updated 9 years ago
- evaluation code for SKU110K dataset☆11Dec 28, 2019Updated 6 years ago
- [CVPR 2024 Highlight] OpenBias: Open-set Bias Detection in Text-to-Image Generative Models☆26Feb 13, 2025Updated last year
- ATS for NeurIPS 2021☆25Nov 4, 2021Updated 4 years ago
- [ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing☆28May 11, 2026Updated 3 months ago
- [Pattern Recognition 2025 🌟]Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation☆10Jun 12, 2024Updated 2 years ago