A Guided Reinforcement Learning framework enhancing MLLM reasoning via process-level verification and collaborative rollout strategies.
☆49May 4, 2026Updated 2 months ago
Alternatives and similar repositories for Guided-GRPO
Users that are interested in Guided-GRPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flow-Modulated Scoring for Semantic-Aware Knowledge Graph Completion.☆18Mar 25, 2026Updated 4 months ago
- [ACL 2026 Findings] Benchmark for large-scale multi-document analysis☆15Jul 14, 2026Updated 2 weeks ago
- 🤖Auto Tutor: Batch-contact prospective advisors☆103Jan 18, 2026Updated 6 months ago
- A curated guide to reasoning-enhancement methods for Multimodal Large Language Models (MLLMs), including dataset construction, training s…☆30Apr 20, 2025Updated last year
- ☆19Oct 28, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official implementation of “Towards Cross-View Point Correspondence in Vision-Language Models”.☆15Dec 24, 2025Updated 7 months ago
- 📐 [CVPR 2026] GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models☆18Apr 1, 2026Updated 3 months ago
- Use Hermes Agent as the control plane for local coding agents like Codex, Kimi Code, Claude Code, OpenCode, and Gemini CLI.☆25Jul 22, 2026Updated last week
- 🪐 Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback☆23Jan 29, 2026Updated 6 months ago
- ☆24Jul 8, 2023Updated 3 years ago
- [ACL 2024] Predicting the Unpredictable: Uncertainty-Aware Reasoning over Temporal Knowledge Graphs via Diffusion Process☆21Oct 7, 2024Updated last year
- DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles☆31Mar 8, 2026Updated 4 months ago
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions [TMLR2025]☆35Jan 13, 2026Updated 6 months ago
- [AAAI 2025 oral] Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit☆19Apr 19, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models☆26Apr 16, 2025Updated last year
- [arxiv 2025] SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot☆22Oct 8, 2025Updated 9 months ago
- ☆14Dec 12, 2024Updated last year
- [ICML 2025] Official implementation of the paper "SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling". …☆21Nov 17, 2025Updated 8 months ago
- Samaya AI's FrontierFinance Benchmark Grader☆18Jul 16, 2026Updated last week
- 夏令营截止日期DDL静态网页☆451Updated this week
- LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs☆41Apr 2, 2026Updated 3 months ago
- Providing the answer to "How to do patching on all available SAEs on GPT-2?". It is an official repository of the implementation of the p…☆13Jan 26, 2025Updated last year
- ☆36Apr 13, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Parallel Continuous Chain-of-Thought with Jacobi Iteration. Accepted to EMNLP 2025.☆23Mar 29, 2026Updated 4 months ago
- ☆26Apr 30, 2026Updated 2 months ago
- A compact high-signal benchmark for evaluating frontier agents☆21Updated this week
- [AAAI-2025] The offical code for SiTo (Similarity-based Token Pruning for Stable Diffusion Models)☆46Jun 2, 2025Updated last year