[ICLR 2026] LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards.
☆19Mar 16, 2026Updated 5 months ago
Alternatives and similar repositories for LongRLVR
Users that are interested in LongRLVR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fine-Tuning Pre-trained Transformers into Decaying Fast Weights☆20Oct 9, 2022Updated 3 years ago
- Long Context Research☆37Aug 31, 2026Updated 2 weeks ago
- [ICML 2025 Spotlight] RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding☆22Mar 2, 2025Updated last year
- Official implementation of MATPO: Multi-Agent Tool-Integrated Policy Optimization.☆82Oct 31, 2025Updated 10 months ago
- This repository includes the implementation and results of the paper "ChatGPT is fun, but it is not funny! Humor is still challenging Lar…☆13Jul 13, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NeurIPS 25]SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset☆20Sep 19, 2025Updated 11 months ago
- ☆11Aug 10, 2024Updated 2 years ago
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents☆45Apr 13, 2026Updated 5 months ago
- [ACL26 Findings] LiCoMemory: Lightweight and Cognitive Agentic Memory for Efficient Long-Term Reasoning☆51Jan 6, 2026Updated 8 months ago
- The official implementation of LIFT: Improving Long Context Understanding of Large Language Models through Long Input Fine-Tuning☆15Mar 14, 2025Updated last year
- ☆14Feb 28, 2024Updated 2 years ago
- [ACL 2026] VGPO: Visually-Guided Policy Optimization for Multimodal Reasoning☆34Apr 14, 2026Updated 5 months ago
- End-to-end Task-oriented Dialog System with Hybrid Knowledge Management☆17Sep 25, 2021Updated 4 years ago
- Code and data for EMNLP2019 Paper "Uncover the Ground-Truth Relations in Distant Supervision: A Neural Expectation-Maximization Framework…☆10May 24, 2020Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆14Feb 26, 2024Updated 2 years ago
- ☆17Jul 31, 2025Updated last year
- ☆16Aug 11, 2025Updated last year
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- ☆10Nov 14, 2021Updated 4 years ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 8 months ago
- lightsmile个人的用于爬取网络公开语料数据的mini通用爬虫框架。☆13Sep 30, 2020Updated 5 years ago
- knrm文本相似度☆10Aug 1, 2020Updated 6 years ago
- 2022 USTC 011705 (OSH) Course Project of Runikraft Group☆13Jul 22, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- 我在ACM校队里学的算法,其中包含图论的大部分算法☆10Nov 19, 2017Updated 8 years ago
- code for "Fine-grained Entity Typing via Label Reasoning" EMNLP2021☆13May 27, 2022Updated 4 years ago
- HiPRAG (Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation) is a reinforcement learning method designed fo…☆27Oct 10, 2025Updated 11 months ago
- ☆12Aug 31, 2021Updated 5 years ago
- ☆10Jun 13, 2020Updated 6 years ago
- Code for paper: Weakly- and Semi-supervised Evidence Extraction☆15Apr 12, 2021Updated 5 years ago
- Official code for "Flatten Graphs as Sequences: Transformers are scalable graph generators" (NeurIPS 2025)☆18Oct 17, 2025Updated 10 months ago
- Agent-Omit: Training Efficient LLM Agents for Adaptive Thought and Observation Omission via Reinforcement Learning☆34May 11, 2026Updated 4 months ago
- Codebase for multilingual neural machine translation☆13Nov 24, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ACL2026 Findings] GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning☆84Jun 23, 2025Updated last year
- Official implementation of 'RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training', accepted by ICLR 2026☆19Oct 15, 2025Updated 11 months ago
- ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model☆15Apr 7, 2026Updated 5 months ago
- ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization☆98May 22, 2025Updated last year
- Make math learning simpler, starting with Nano Math plus , too!☆45Jul 16, 2026Updated last month
- Official code of ConfTuner: Training Large Language Models to Express Their Confidence Verbally☆27Sep 26, 2025Updated 11 months ago
- The official Pytorch implementation of "UniKER: A Unified Framework for Combining Embedding and Definite Horn Rule Reasoning for Knowledg…☆15May 16, 2022Updated 4 years ago