[ICLR'26] SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
☆18Mar 26, 2026Updated 5 months ago
Alternatives and similar repositories for SketchThinker-R1
Users that are interested in SketchThinker-R1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2026] Code for "The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts"☆21Jun 13, 2026Updated 3 months ago
- ☆16Jan 13, 2024Updated 2 years ago
- SEED Dataset☆29Jun 3, 2025Updated last year
- Offical repo for ECCV 2024: Depth-Aware Blind Image Decomposition for Real-World Weather Recovery☆13Mar 7, 2024Updated 2 years ago
- 🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation☆25Jul 11, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official repo of "Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark"☆21Jun 5, 2025Updated last year
- This is the official implementation of "Transferring to Real-World Layouts: A Depth-aware Framework for Scene Adaptation" (Accepted at AC…☆13Aug 24, 2024Updated 2 years ago
- ✨A curated list of papers on the uncertainty in multi-modal large language model (MLLM).☆59Apr 2, 2025Updated last year
- ✨ Official code for our paper: "Uncertainty-o: One Model-agnostic Framework for Unveiling Epistemic Uncertainty in Large Multimodal Model…☆21Mar 13, 2025Updated last year
- 📖Curated list about reasoning abilitiy of MLLM, including OpenAI o1, OpenAI o3-mini, and Slow-Thinking.☆13Feb 7, 2025Updated last year
- ACM MM Workshop on UAVs in Multimedia: Capturing the World from a New Perspective (UAVM 2023)☆13Jul 4, 2026Updated 2 months ago
- The official code of "Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search"☆38Jul 25, 2026Updated last month
- You Only Condense Once: Two Rules for Pruning Condensed Datasets (NeurIPS 2023)☆17Jul 30, 2026Updated last month
- An offical repo for ECCV 2024 Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching☆120Jul 7, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Workshop on UAVs in Multimedia: Capturing the World from a New Perspective. Reza Zhu's Solution: MBEG☆11May 17, 2024Updated 2 years ago
- [ICLR 2020] Haotao Wang, Tianlong Chen, Zhangyang Wang, Kede Ma, "I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifie…☆20Dec 30, 2021Updated 4 years ago
- Codes, data, and baselines for CIKM 2023 Long Paper "Dual Intents Graph Modeling for User-centric Group Discovery"☆17Oct 22, 2023Updated 2 years ago
- Set of scripts and instructions for sub-selecting and formatting raw data exported by the Canfield ISIC2024 Tile Export Tool. The resulti…☆12May 8, 2024Updated 2 years ago
- 🔎Official code for our paper: "VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation".☆56Mar 18, 2025Updated last year
- Graph-Backed Generative Brick Assembly☆41Aug 19, 2026Updated last month
- ☆15Jan 6, 2025Updated last year
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 9 months ago
- The official code of "CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval"☆15Sep 19, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs☆18Feb 3, 2026Updated 7 months ago
- [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting R…☆49Mar 1, 2026Updated 6 months ago
- 复旦研究生抢课脚本☆12Feb 14, 2022Updated 4 years ago
- [ICLR 2026 🔥 ] Official implementation of "UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing"☆152Jan 26, 2026Updated 7 months ago
- Codes and data for CIKM 2022 paper "RuDi: Explaining Behavior Sequence Models by Automatic Statistics Generation and Rule Distillation"☆12Aug 16, 2022Updated 4 years ago
- ☆43Jun 10, 2025Updated last year
- [CVPR 2024] Official repository of ST_GT☆10Sep 15, 2024Updated 2 years ago
- ☆23Nov 24, 2022Updated 3 years ago
- ICME2022 Special Session “Beyond Accuracy: Responsible, Responsive, and Robust Multimedia Retrieval ”☆12Jun 3, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆10Oct 18, 2023Updated 2 years ago
- Cornell Tech CS5670 Introduction to Computer Vision Projects Repo☆13Nov 22, 2022Updated 3 years ago
- [Pattern Recognition'24] Pytorch implementation of Multiple-environment Self-adaptive Network for Aerial-view Geo-localization https://a…☆47Jul 6, 2026Updated 2 months ago
- 地图足迹故事,微信小程序☆10May 5, 2022Updated 4 years ago
- [SIGKDD 2024] Rethinking Fair Graph Neural Networks from Re-balancing☆10Jul 15, 2024Updated 2 years ago
- A curated list of papers on graph transfer learning (GTL).☆19Oct 23, 2023Updated 2 years ago
- ☆22Nov 5, 2024Updated last year