[ICLR 2026] LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards.
☆19Mar 16, 2026Updated 4 months ago
Alternatives and similar repositories for LongRLVR
Users that are interested in LongRLVR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- Fine-Tuning Pre-trained Transformers into Decaying Fast Weights☆20Oct 9, 2022Updated 3 years ago
- [ICML 2025 Spotlight] RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding☆22Mar 2, 2025Updated last year
- Official implementation of MATPO: Multi-Agent Tool-Integrated Policy Optimization.☆82Oct 31, 2025Updated 9 months ago
- This repository includes the implementation and results of the paper "ChatGPT is fun, but it is not funny! Humor is still challenging Lar…☆13Jul 13, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [NeurIPS 25]SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset☆21Sep 19, 2025Updated 10 months ago
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents☆44Apr 13, 2026Updated 3 months ago
- The official implementation of LIFT: Improving Long Context Understanding of Large Language Models through Long Input Fine-Tuning☆15Mar 14, 2025Updated last year
- ☆14Feb 28, 2024Updated 2 years ago
- End-to-end Task-oriented Dialog System with Hybrid Knowledge Management☆17Sep 25, 2021Updated 4 years ago
- Code and data for EMNLP2019 Paper "Uncover the Ground-Truth Relations in Distant Supervision: A Neural Expectation-Maximization Framework…☆10May 24, 2020Updated 6 years ago
- SDPG: Self-Distilled Policy Gradient☆50Jun 15, 2026Updated last month
- ☆18Jul 31, 2025Updated last year
- Code for paper "ProgGen: Generating Named Entity Recognition Datasets Step-by-step with Self-Reflexive Large Language Models"☆17Mar 29, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- ☆10Nov 14, 2021Updated 4 years ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆25Dec 30, 2025Updated 7 months ago
- [EMNLP 2023] Once Upon a *Time* in *Graph*: Relative-Time Pretraining for Complex Temporal Reasoning☆17Oct 31, 2023Updated 2 years ago
- lightsmile个人的用于爬取网络公开语料数据的mini通用爬虫框架。☆13Sep 30, 2020Updated 5 years ago
- 污染源在线自动监控(监测)系统数据传输标准,数据包模拟发送程序、检测程序。支持水、气、TVOC☆12Sep 14, 2018Updated 7 years ago
- knrm文本相似度☆10Aug 1, 2020Updated 6 years ago
- 2022 USTC 011705 (OSH) Course Project of Runikraft Group☆13Jul 22, 2022Updated 4 years ago
- A comprehensive and efficient long-context model evaluation framework☆31Feb 25, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- HiPRAG (Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation) is a reinforcement learning method designed fo…☆26Oct 10, 2025Updated 9 months ago
- code for "Fine-grained Entity Typing via Label Reasoning" EMNLP2021☆13May 27, 2022Updated 4 years ago
- code for paper Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering☆14Aug 13, 2024Updated last year
- ☆15Aug 12, 2022Updated 3 years ago
- ☆25Jan 22, 2024Updated 2 years ago
- ☆12Aug 31, 2021Updated 4 years ago
- LLM model connection LangChain RAG Connection to Streamlit Web☆14Oct 22, 2023Updated 2 years ago
- ☆10Jun 13, 2020Updated 6 years ago
- Code for paper: Weakly- and Semi-supervised Evidence Extraction☆15Apr 12, 2021Updated 5 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆15Jan 9, 2018Updated 8 years ago
- Image Classification Tutorial: ConvNext--> 98.8% on CIFAR10 + 92.4% on CIFAR100; ResNet18 -- 95.6% on CIFAR10 + 79.1% on CIFAR100☆15Jun 2, 2025Updated last year
- [ACL2026 Findings] GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning☆83Jun 23, 2025Updated last year
- Official repo for "PAPO: Perception-Aware Policy Optimization for Multimodal Reasoning"☆153Feb 4, 2026Updated 6 months ago
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆22May 6, 2026Updated 3 months ago
- Official implementation of 'RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training', accepted by ICLR 2026☆18Oct 15, 2025Updated 9 months ago
- ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model☆15Apr 7, 2026Updated 3 months ago