[ACL 2026] VGPO: Visually-Guided Policy Optimization for Multimodal Reasoning
☆36Apr 14, 2026Updated 5 months ago
Alternatives and similar repositories for VGPO
Users that are interested in VGPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GaitParsing: Human Semantic Parsing for Gait Recognition (IEEE TMM)☆13May 20, 2024Updated 2 years ago
- [AAAI 2024] QAGait: Revisit Gait Recognition From a Quality Perspective☆23Aug 26, 2024Updated 2 years ago
- ☆24Feb 2, 2026Updated 8 months ago
- Official implementation of "Figure It Out: Improve the Frontier of Reasoning with Active Visual Thinking"☆17Jan 13, 2026Updated 8 months ago
- [EMNLP’ 25] Official code for "HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation…☆37Nov 3, 2025Updated 11 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆62Feb 9, 2026Updated 8 months ago
- [ICCV2025] UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement☆29Apr 25, 2026Updated 5 months ago
- [ICML 2026] Heima☆77May 20, 2026Updated 4 months ago
- [ICLR2026] Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models☆151Jan 30, 2026Updated 8 months ago
- [ICML 2026] The official implementation of paper "Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation…☆95Updated this week
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 9 months ago
- [ICLR 2026] FASA: FREQUENCY-AWARE SPARSE ATTENTION☆22Aug 18, 2026Updated last month
- MemOCR: an OCR-driven visual memory agent.☆35May 17, 2026Updated 4 months ago
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆18Oct 25, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 8 months ago
- Syphus: Automatic Instruction-Response Generation Pipeline☆14Dec 14, 2023Updated 2 years ago
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models☆19Sep 5, 2026Updated last month
- [ICLR 2026] LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards.☆19Mar 16, 2026Updated 6 months ago
- ☆22May 4, 2026Updated 5 months ago
- [ACL-26 (main)] From Verbatim to Gist Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video A…☆43Apr 19, 2026Updated 5 months ago
- [2026 CVPR]Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation☆109Apr 15, 2026Updated 5 months ago
- [ICCV25] LD-RPS☆49Sep 7, 2026Updated last month
- [CVPR 2026] Elucidating the SNR-t Bias of Diffusion Probabilistic Models☆121Apr 20, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A comprehensive benchmark specifically designed to evaluate the interactive response capabilities of world models in 4D settings.☆108Mar 24, 2026Updated 6 months ago
- ☆29Aug 8, 2025Updated last year
- [NeurIPS 2025 Spotlight] Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning☆55Apr 16, 2026Updated 5 months ago
- Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"☆296May 28, 2026Updated 4 months ago
- [ICLR 26] Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow☆47Oct 3, 2025Updated last year
- ☆21Aug 30, 2026Updated last month
- ☆24Jun 16, 2026Updated 3 months ago
- Code of Self-Distilled RLVR - RLSD☆77May 19, 2026Updated 4 months ago
- [ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models☆72May 15, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- code of cvpr26 paper Symphony☆17Apr 7, 2026Updated 6 months ago
- [ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"☆91Apr 3, 2026Updated 6 months ago
- TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation☆129May 30, 2026Updated 4 months ago
- [ICLR 2026] Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation☆129May 17, 2026Updated 4 months ago
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆20Jun 2, 2026Updated 4 months ago
- ☆96Feb 5, 2026Updated 8 months ago
- Official implementation of High-Fidelity Zero-Shot Texture Anomaly Localization Using Feature Correspondence Analysis.☆12Dec 18, 2023Updated 2 years ago