[CVPR 2026] GThinker, Reasoning MLLM, Visual Cues, Visual Rethinking
☆18Mar 9, 2026Updated 6 months ago
Alternatives and similar repositories for GThinker
Users that are interested in GThinker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner☆24Aug 7, 2025Updated last year
- Recursive Abstractive Processing for Tree-Organized Retrieval☆10May 30, 2024Updated 2 years ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 4 months ago
- ☆34Sep 19, 2025Updated last year
- Official repo of Griffon series including v1(ECCV 2024), v2(ICCV 2025), G, and R, and also the RL tool Vision-R1(CVPR 2026).☆252Apr 17, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Grounded Visual Token Sampling (GroundVTS), a Vid-LLM architecture designed to enhance VTG performance through adaptive and efficient vis…☆18Jun 12, 2026Updated 3 months ago
- Training codebase for K2-V2☆22Dec 17, 2025Updated 9 months ago
- Official Repo for CVPR 2025 Paper -- DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos☆17Mar 16, 2026Updated 6 months ago
- [ACM MM 2025 🔥🔥 ] MIRA: A first-of-its-kind medical RAG framework that fuses image features and retrieved knowledge with dynamic contex…☆23Aug 28, 2025Updated last year
- Weakly-supervised road-lane markings detection for autonomous driving, mitigating the lack of training data☆14Oct 8, 2024Updated last year
- KDD25: The source code of our paper "UoMo: A Universal Model of Mobile Traffic Forecasting for Wireless Network Optimization"☆16Aug 14, 2025Updated last year
- Our 2nd-gen LMM☆34May 22, 2024Updated 2 years ago
- (CVPR 26 Findings) Official implementation of the paper "Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-…☆34Apr 7, 2026Updated 5 months ago
- [Official, NeurIPS 2025] TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs.☆26Jun 8, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- We introduce DreamPRM-1.5, an instance-reweighted framework that adaptively adjusts the importance of each training example via bi-level …☆16Nov 13, 2025Updated 10 months ago
- The official repo for "Unified Domain Adaptive Semantic Segmentation" (IEEE TPAMI 2025)☆33Aug 14, 2025Updated last year
- An interactive thinking and deep reasoning model. It provides a cognitive reasoning paradigm for complex multi-hop problems.☆89Nov 14, 2025Updated 10 months ago
- Self-Teaching Notes on Gradient Leakage Attacks against GPT-2 models.☆14Mar 18, 2024Updated 2 years ago
- Multimodal Federated Learning on IoT Data☆11Dec 17, 2023Updated 2 years ago
- [NeurIPS 2025] Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing☆99Jul 27, 2025Updated last year
- EVA: Efficient Reinforcement Learning for End-to-End Video Agent☆26May 6, 2026Updated 4 months ago
- ☆16Aug 19, 2026Updated last month
- [WACV 2024] Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining, WACV 2024☆13Jan 3, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Qwen-WisdomVast is a large model trained on 1 million high-quality Chinese multi-turn SFT data, 200,000 English multi-turn SFT data, and …☆17Apr 12, 2024Updated 2 years ago
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement …☆44Aug 6, 2025Updated last year
- Advanced Embodied Intelligence Brain Model☆37Nov 5, 2025Updated 10 months ago
- [ICML 2026] Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval☆25Sep 9, 2026Updated last week
- [CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice☆89Feb 27, 2026Updated 6 months ago
- [ICML 2026] VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding☆29Aug 10, 2026Updated last month
- [ECCV 2026] Official implementation of CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding.☆52Sep 15, 2025Updated last year
- Official repo of Promoting Efficient Reasoning with Verifiable Stepwise Reward☆16Sep 9, 2025Updated last year
- The official repository of "R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Integration"☆141Sep 4, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning☆15Jun 2, 2026Updated 3 months ago
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsing☆28Jan 19, 2026Updated 8 months ago
- Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption☆38May 21, 2025Updated last year
- View planning with multi-turn VLM agents: ViewSuite 6-DoF benchmark on real ScanNet scenes + iterative RL-SFT training☆25Sep 6, 2026Updated last week
- Artifact evaluation of MobiSys25 SynCheck☆20Mar 24, 2025Updated last year
- Official code of *Towards Event-oriented Long Video Understanding*☆12Jul 26, 2024Updated 2 years ago
- ☆35Feb 12, 2026Updated 7 months ago