[ICLR'26] SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
☆17Mar 26, 2026Updated 4 months ago
Alternatives and similar repositories for SketchThinker-R1
Users that are interested in SketchThinker-R1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repo for Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting☆17Mar 31, 2026Updated 4 months ago
- [CVPR 2026] Code for "The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts"☆20Jun 13, 2026Updated last month
- ☆16Jan 13, 2024Updated 2 years ago
- SEED Dataset☆29Jun 3, 2025Updated last year
- [ECCV'24] Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene.☆40Sep 3, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- 🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation☆23Jul 11, 2026Updated last month
- Progressive Text-to-3D Generation for Automatic 3D Prototyping (ACM TOMM)☆54Mar 14, 2026Updated 4 months ago
- The official repo of "Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark"☆21Jun 5, 2025Updated last year
- 📖Curated list about reasoning abilitiy of MLLM, including OpenAI o1, OpenAI o3-mini, and Slow-Thinking.☆13Feb 7, 2025Updated last year
- ACM MM Workshop on UAVs in Multimedia: Capturing the World from a New Perspective (UAVM 2023)☆13Jul 4, 2026Updated last month
- Official Repo for ICCV25-Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization☆45Jul 6, 2026Updated last month
- An offical repo for ECCV 2024 Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching☆119Jul 7, 2026Updated last month
- ICLR‘24 Offical Implementation of Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization☆74Jan 30, 2024Updated 2 years ago
- Official implementation for paper "Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe"☆39May 12, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR 2020] Haotao Wang, Tianlong Chen, Zhangyang Wang, Kede Ma, "I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifie…☆20Dec 30, 2021Updated 4 years ago
- [ICCV'25] "Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection".☆26Jan 12, 2026Updated 6 months ago
- [npj AI] 3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective Single-View 3D Reconstruction☆90Jul 6, 2026Updated last month
- Codes, data, and baselines for CIKM 2023 Long Paper "Dual Intents Graph Modeling for User-centric Group Discovery"☆17Oct 22, 2023Updated 2 years ago
- ☆11Mar 13, 2017Updated 9 years ago
- 🔎Official code for our paper: "VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation".☆56Mar 18, 2025Updated last year
- Official Implementation of PiPa: Pixel- and Patch-wise Self-supervised Learning for Domain Adaptative Semantic Segmentation☆100Jul 23, 2024Updated 2 years ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 7 months ago
- Inpainting using RunwayML's stable-diffusion-inpainting checkpoint☆19Jan 15, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The official code of "CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval"☆15Sep 19, 2024Updated last year
- [ECCV 2026] VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs☆17Feb 3, 2026Updated 6 months ago
- ☆21Oct 10, 2020Updated 5 years ago
- [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting R…☆43Mar 1, 2026Updated 5 months ago
- 复旦研究生抢课脚本☆10Feb 14, 2022Updated 4 years ago
- [ICLR 2026 🔥 ] Official implementation of "UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing"☆151Jan 26, 2026Updated 6 months ago
- UM CIS PhD Qualifying Examination Resources - A curated collection of study notes, past materials, and preparation guides for the Qualify…☆19May 30, 2026Updated 2 months ago
- ☆44Jun 10, 2025Updated last year
- [CVPR 2024] Official repository of ST_GT☆10Sep 15, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆23Nov 24, 2022Updated 3 years ago
- ICME2022 Special Session “Beyond Accuracy: Responsible, Responsive, and Robust Multimedia Retrieval ”☆12Jun 3, 2024Updated 2 years ago
- ☆10Oct 18, 2023Updated 2 years ago
- ResearchDoom fork of the Chocolate Doom engine.☆16Updated this week
- ☆22Nov 5, 2024Updated last year
- [CVPR 2025] This repository is intended to store the code and data for ASAP (Advancing Semantic Alignment Promotes Multi-Modal Manipulati…☆21Jun 18, 2025Updated last year
- TVBench: Redesigning Video-Language Evaluation☆15Jun 9, 2025Updated last year