[NeurIPS 2025 Spotlight] Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
☆55Apr 16, 2026Updated 5 months ago
Alternatives and similar repositories for FAST
Users that are interested in FAST are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [AAAI 2025] Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback☆36Dec 16, 2025Updated 9 months ago
- [ACL 2026] G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance☆16Sep 15, 2026Updated last week
- This repo holds the official code and data for "Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with H…☆15May 21, 2024Updated 2 years ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation☆112Sep 18, 2025Updated last year
- Official code for the paper, "TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible Adapter".☆16Jun 20, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR 2026] Rectifying LLM Thought From Lens of Optimization☆14Dec 5, 2025Updated 9 months ago
- CVPR25☆28Jul 2, 2025Updated last year
- The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]☆192Jun 5, 2025Updated last year
- (ICML 2025) Rethinking Chain-of-Thought from the Perspective of Self-Training☆13Feb 15, 2025Updated last year
- [ICML 2026] Heima☆77May 20, 2026Updated 4 months ago
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆56Oct 9, 2025Updated 11 months ago
- [ACL 2026] VGPO: Visually-Guided Policy Optimization for Multimodal Reasoning☆35Apr 14, 2026Updated 5 months ago
- [ICCV 2025] Boosting MLLM Reasoning with Text-Debiased Hint-GRPO☆48Jul 1, 2025Updated last year
- AAAI '25. Retrieval-Augmented Multimodal Social Media Popularity Prediction☆26Sep 15, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Syphus: Automatic Instruction-Response Generation Pipeline☆14Dec 14, 2023Updated 2 years ago
- Code for "Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning [EMNLP 2025 Findings]"☆18Aug 27, 2025Updated last year
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning☆56Jul 23, 2025Updated last year
- [NeurIPS25] Official Implementation (Pytorch) of "DeepVideo-R1"☆38Feb 22, 2026Updated 7 months ago
- ☆27Mar 26, 2026Updated 5 months ago
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward☆95Aug 8, 2025Updated last year
- Official Repository of Personalized Visual Instruct Tuning☆34Mar 6, 2025Updated last year
- ☆16Aug 18, 2025Updated last year
- ☆27Apr 3, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A hierarchical multi-agent framework for exhaustive cross-document question answering.☆21Mar 14, 2026Updated 6 months ago
- [TMLR 25] SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models☆147Oct 10, 2025Updated 11 months ago
- [CVPR 24] This is official implication for our paper: ''CroSel: Cross Selection of Confident Pseudo Labels for Partial-Label Learning''.☆15Apr 27, 2025Updated last year
- The IOS app Right to Record, which is an app that records and uploads at the same time, this can protect your video footages so that even…☆76Dec 2, 2025Updated 9 months ago
- ☆29Aug 8, 2025Updated last year
- Code for "SePPO: Semi-Policy Preference Optimization for Diffusion Alignment."☆18Oct 7, 2024Updated last year
- AIFlow is an AI agentic framework designed to scale digital AI agents on BNB Chain.☆242Feb 28, 2025Updated last year
- Official code repository of Shuffle-R1☆26Feb 23, 2026Updated 7 months ago
- Official code for the paper: DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models☆24Jan 6, 2026Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR 2026] MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources☆217Sep 26, 2025Updated 11 months ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆25Jun 26, 2026Updated 2 months ago
- 【ICLR 2026】 Official Repo for Paper ‘’OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis‘’☆22Mar 4, 2026Updated 6 months ago
- ☆14Nov 26, 2025Updated 9 months ago
- [ACL '26] source code for the paper: "Long-Chain Reasoning Distillation via Adaptive Prefix Alignment"☆17Jan 21, 2026Updated 8 months ago
- Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.☆3,171Jul 29, 2026Updated last month
- Official Implementation of Trajectory-Refined Distillation☆37Jun 9, 2026Updated 3 months ago