[NeurIPS 2025 Spotlight] Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning
☆89Sep 19, 2025Updated 10 months ago
Alternatives and similar repositories for CLS-RL
Users that are interested in CLS-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of "Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck"☆16Nov 10, 2025Updated 8 months ago
- ☆15Sep 14, 2023Updated 2 years ago
- [EMNLP 2024] Implementation of vision-language model fine-tuning via simple parameter-efficient modification☆19Nov 24, 2024Updated last year
- [TMLR 25] SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models☆148Oct 10, 2025Updated 9 months ago
- Explore the Multimodal “Aha Moment” on 2B Model☆624Mar 18, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]☆189Jun 5, 2025Updated last year
- [ICLR 2025 Spotlight] Realistic Evaluation of Deep Partial-Label Learning Algorithms☆15Feb 2, 2025Updated last year
- [ICLR'25] Official repository of paper titled "Tree of Attributes Prompt Learning for Vision-Language Models".☆20Oct 15, 2025Updated 9 months ago
- ✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning☆291May 9, 2025Updated last year
- Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.☆848May 14, 2025Updated last year
- Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’☆2,262Oct 29, 2025Updated 8 months ago
- [CVPR2025] Official implementation of RAM☆29Nov 4, 2025Updated 8 months ago
- Consistent Prompting for Rehearsal-Free Continual Learning [CVPR2024]☆34Jun 12, 2025Updated last year
- Official repo for the TMLR paper "Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners"☆29Apr 27, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The official code for "TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning" | [AAAI2025]☆53Mar 13, 2025Updated last year
- [CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"☆84May 12, 2026Updated 2 months ago
- Official implementation of the paper “Endowing Vision-Language Models with System 2 Thinking for Fine-Grained Visual Recognition,” AAAI 2…☆44Jan 30, 2026Updated 5 months ago
- [ICLR 2025] "Noisy Test-Time Adaptation in Vision-Language Models"☆13Feb 22, 2025Updated last year
- Code for BYOP [CVPR 2023]☆11Sep 25, 2023Updated 2 years ago
- my attempt at implementing the DiffEdit paper (WIP)☆16Oct 30, 2022Updated 3 years ago
- [MTI-LLM@NeurIPS 2025] Official implementation of "PyVision: Agentic Vision with Dynamic Tooling."☆162Jul 22, 2025Updated last year
- VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models☆79Jul 13, 2024Updated 2 years ago
- Codebase for VidHal: Benchmarking Hallucinations in Vision LLMs☆14Apr 23, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NeurIPS 2023] Official repository for "Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models"☆11Jun 18, 2024Updated 2 years ago
- [Blog 1] Recording a bug of grpo_trainer in some R1 projects☆23Feb 23, 2025Updated last year
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation☆112Sep 18, 2025Updated 10 months ago
- Witness the aha moment of VLM with less than $3.☆4,065May 19, 2025Updated last year
- TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning☆116Dec 24, 2025Updated 7 months ago
- Code for Label Propagation for Zero-shot Classification with Vision-Language Models (CVPR2024)☆45Jul 23, 2024Updated 2 years ago
- [ICCV 2025] MMReason, MLLMs, step by step, reasoning benchmark, AGI☆15Apr 25, 2026Updated 2 months ago
- Code implementation of our ICCV 2025 paper: On Large Multimodal Models as Open-World Image Classifiers☆27Dec 4, 2025Updated 7 months ago
- (CVPR 2025) PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction☆151Mar 6, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆105Jun 10, 2025Updated last year
- ☆15Jan 14, 2026Updated 6 months ago
- OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.☆399Jun 1, 2025Updated last year
- ☆132Jul 22, 2025Updated last year
- ☆12Feb 27, 2025Updated last year
- ☆68Apr 4, 2026Updated 3 months ago
- [NeurIPS 2025] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models☆82May 31, 2025Updated last year