[ICLR 2026] Quantile Advantage Estimation for Entropy-Safe Reasoning
☆29Oct 14, 2025Updated 9 months ago
Alternatives and similar repositories for QAE
Users that are interested in QAE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆22Mar 11, 2026Updated 5 months ago
- [ICML 2025] Official code of "DAMA: Data- and Model-aware Alignment of Multi-modal LLMs"☆16May 24, 2025Updated last year
- Source code of "Training Free Graph Neural Networks for Graph Matching"☆12Jul 9, 2022Updated 4 years ago
- This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.☆18Jul 14, 2026Updated 3 weeks ago
- ☆27Jan 20, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [NeurIPS 2024] Official code of $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$☆51Oct 23, 2024Updated last year
- NeurIPS 2025: Discriminative Constrained Optimization for Reinforcing Large Reasoning Models☆53Mar 14, 2026Updated 4 months ago
- Official implementation for "ALI-Agent: Assessing LLMs'Alignment with Human Values via Agent-based Evaluation"☆21Jan 31, 2026Updated 6 months ago
- PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations☆13Apr 21, 2024Updated 2 years ago
- [EMNLP 2023] ReLM: Leveraging Language Models for Enhanced Chemical Reaction Prediction.☆22Jan 28, 2024Updated 2 years ago
- Accompanying code for our NeurIPS 2019 paper☆11Nov 7, 2019Updated 6 years ago
- Archer2.0 evolves from its predecessor by introducing ASPO, which overcomes fundamental PPO-Clip limitations to prevent premature converg…☆31Oct 10, 2025Updated 10 months ago
- AnyEdit: Edit Any Knowledge Encoded in Language Models, ICML 2025☆49Nov 6, 2025Updated 9 months ago
- ☆41Nov 20, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆25Oct 23, 2025Updated 9 months ago
- ☆25Nov 16, 2023Updated 2 years ago
- ☆13Jan 22, 2025Updated last year
- [ICLR 2025] Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist☆34Oct 23, 2024Updated last year
- Official implementation for Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos☆16May 23, 2023Updated 3 years ago
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents☆44Apr 13, 2026Updated 3 months ago
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement …☆44Aug 6, 2025Updated last year
- ☆13Feb 25, 2025Updated last year
- [NeurIPS2023] Official code of "Understanding Contrastive Learning via Distributionally Robust Optimization"☆40Oct 18, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NeurIPS 2023] "Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules"☆41Mar 16, 2024Updated 2 years ago
- A Deep Reinforcement Learning Strategy and Framework for Floating Waste Capture☆13Mar 13, 2025Updated last year
- ☆19Sep 16, 2025Updated 10 months ago
- LeetCode Training and Evaluation Dataset☆54Apr 22, 2025Updated last year
- [WWW 2023] Official code of "Adap-$\tau$: Adaptively Modulating Embedding Magnitude for Recommendation"☆29Jan 4, 2024Updated 2 years ago
- repository for "Exploiting Proximity-Aware Tasks for Embodied Social Navigation" paper code☆12Nov 16, 2023Updated 2 years ago
- ☆10Apr 8, 2024Updated 2 years ago
- ☆14Jun 11, 2024Updated 2 years ago
- [ICLR 2026] Geometric-Mean Policy Optimization☆104Jan 26, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The repository of CLEME (EMNLP 2023) and CLEME2.0 (ACL 2025)☆12May 17, 2025Updated last year
- source code of (quasi-)Givens Orthogonal Fine Tuning integrated to peft lib☆16Mar 13, 2025Updated last year
- Code for EMNLP2023 paper "MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter".☆12Dec 27, 2023Updated 2 years ago
- Research Code for preprint "Optimizing Test-Time Compute via Meta Reinforcement Finetuning".☆120Jun 23, 2026Updated last month
- [ACL'23 Findings] This is the code repo for our ACL'23 Findings paper "ReGen: Zero-Shot Text Classification via Training Data Generation …☆24Sep 8, 2023Updated 2 years ago
- Official implementation of ECCV24 paper: POA☆24Aug 8, 2024Updated 2 years ago
- This is a pip package implementing Reinforcement Learning algorithms in non-stationary environments supported by the OpenAI Gym toolkit.☆16Jun 28, 2024Updated 2 years ago