Rewards as Labels: Revisiting RLVR from a Classification Perspective
β25Jun 26, 2026Updated 3 months ago
Alternatives and similar repositories for REAL
Users that are interested in REAL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- πͺ Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedbackβ26Jan 29, 2026Updated 8 months ago
- A Multimodal Reasoning Agent with Stateful Experiencesβ28Mar 31, 2026Updated 6 months ago
- Sparking "Thinking with Videos" via Reinforcement Learningβ159Oct 30, 2025Updated 11 months ago
- [MICCAI 2024] VLSM-Adapter: Finetuning Vision-Language Segmentation Efficiently with Lightweight Blocksβ29Sep 24, 2026Updated 2 weeks ago
- HyperEyes is a parallel multimodal search agent that fuses visual grounding and retrieval into a single atomic action, enabling concurrenβ¦β76May 23, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- (ICLR 2025) AgentRefine: Enhancing Agent Generalization through Refinement Tuningβ22Nov 22, 2025Updated 10 months ago
- [CVRP 2025] Symbolic Representation for Any-to-Any Generative Tasksβ19Mar 24, 2025Updated last year
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to β¦β75Jan 28, 2026Updated 8 months ago
- [NeurIPS 2026 Pre-to-Post ] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Spaceβ18Apr 16, 2026Updated 5 months ago
- π οΈ DeepAgent: A General Reasoning Agent with Scalable Toolsetsβ53Dec 31, 2025Updated 9 months ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervisionβ107Jul 23, 2026Updated 2 months ago
- [ACL 2024 Findings] Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruningβ13Sep 2, 2024Updated 2 years ago
- Official repository for ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Useβ31Nov 4, 2025Updated 11 months ago
- Core Library of Discrete Distribution Networks (ICLR 2025)β16Oct 12, 2025Updated 11 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- OmniGAIA: Towards Native Omni-Modal AI Agentsβ148Apr 2, 2026Updated 6 months ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Miningβ55Apr 22, 2026Updated 5 months ago
- {DeepL, Google, WMT-Best, davinci-003, turbo, gpt-4} Γ {En-De, En-Cs, En-Ru, En-Zh, De-Fr, En-Ja, Uk-En, Uk-Cs, En-Hr, En-Ha, En-Is}β14Jun 18, 2023Updated 3 years ago
- π΅ Code for Less is More for Long Document Summary Evaluation by LLMs (Wu*, Iso* et al; EACL 2024)β12Feb 22, 2024Updated 2 years ago
- SpyGame: An interactive multi-agent framework to evaluate intelligence with large language models :Dβ15Nov 9, 2023Updated 2 years ago
- This repo contains the code for Late Prompt Tuning.β12Dec 22, 2025Updated 9 months ago
- β20May 1, 2025Updated last year
- [CVPR'26] SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimizationβ23Feb 19, 2026Updated 7 months ago
- β27Jun 10, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Towards Defending against Adversarial Examples via Attack-Invariant Featuresβ12Oct 12, 2023Updated 2 years ago
- Implementation of an LLM prompting pipeline combined with wrappers for auto-decomposing reasoning steps and for search through the reasonβ¦β16May 7, 2024Updated 2 years ago
- β47Apr 9, 2025Updated last year
- β15Jul 16, 2021Updated 5 years ago
- Training Autoregressive Image Generation models via Reinforcement Learningβ53Nov 26, 2025Updated 10 months ago
- Learning adapter weights from task descriptionsβ21Nov 12, 2023Updated 2 years ago
- β27Feb 18, 2024Updated 2 years ago
- NeurIPS 2025: Discriminative Constrained Optimization for Reinforcing Large Reasoning Modelsβ53Mar 14, 2026Updated 6 months ago
- Implementation of the paper "In-context Time Series Predictor" (ICLR 2025)β18Feb 11, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Source Code for "Adapters for Enhanced Modeling of Multilingual Knowledge and Text"β12Oct 28, 2022Updated 3 years ago
- This repository contains the code for the paper βNeuro-Symbolic Query Compilerβ, accepted to the Findings of ACL 2025.β19Oct 20, 2025Updated 11 months ago
- The source code of our ACL paper "A Training-free and Reference-free Summarization Evaluation Metric via Centrality-weighted Relevance anβ¦β14May 6, 2023Updated 3 years ago
- β51Jul 21, 2026Updated 2 months ago
- The implementation for our paper, "Improving Simultaneous Machine Translation with Monolingual Data," accepted to AAAI 2023. πβ12Jul 19, 2023Updated 3 years ago
- Paper writing guide for Zhuang Liu Lab @ Princeton Universityβ38Jun 24, 2026Updated 3 months ago
- β19Jan 3, 2025Updated last year