[ICLR 2026] Quantile Advantage Estimation for Entropy-Safe Reasoning
☆29Oct 14, 2025Updated 9 months ago
Alternatives and similar repositories for QAE
Users that are interested in QAE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Feb 26, 2025Updated last year
- ☆22Mar 11, 2026Updated 4 months ago
- [ICML 2025] Official code of "DAMA: Data- and Model-aware Alignment of Multi-modal LLMs"☆16May 24, 2025Updated last year
- Source code of "Training Free Graph Neural Networks for Graph Matching"☆12Jul 9, 2022Updated 4 years ago
- This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.☆18Jul 14, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆27Jan 20, 2025Updated last year
- [NeurIPS 2024] Official code of $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$☆51Oct 23, 2024Updated last year
- NeurIPS 2025: Discriminative Constrained Optimization for Reinforcing Large Reasoning Models☆53Mar 14, 2026Updated 4 months ago
- [ICLR'25] PiCO: Peer Review in LLMs based on the Consistency Optimization, https://arxiv.org/pdf/2402.01830☆36Feb 16, 2025Updated last year
- Official implementation for "ALI-Agent: Assessing LLMs'Alignment with Human Values via Agent-based Evaluation"☆21Jan 31, 2026Updated 5 months ago
- A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Models☆41Apr 17, 2026Updated 3 months ago
- Minimal and Customizable CC-Style Coding Agent☆131Apr 2, 2026Updated 3 months ago
- PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations☆12Apr 21, 2024Updated 2 years ago
- Accompanying code for our NeurIPS 2019 paper☆11Nov 7, 2019Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Archer2.0 evolves from its predecessor by introducing ASPO, which overcomes fundamental PPO-Clip limitations to prevent premature converg…☆31Oct 10, 2025Updated 9 months ago
- ☆41Nov 20, 2023Updated 2 years ago
- ☆18Mar 30, 2025Updated last year
- ☆25Nov 16, 2023Updated 2 years ago
- Contrastive self-supervised learning using Rényi divergence☆14Oct 21, 2022Updated 3 years ago
- ☆13Jan 22, 2025Updated last year
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents☆42Apr 13, 2026Updated 3 months ago
- The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general per…☆17Aug 28, 2025Updated 10 months ago
- ☆61Apr 9, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement …☆44Aug 6, 2025Updated 11 months ago
- [NeurIPS 2023] "Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules"☆41Mar 16, 2024Updated 2 years ago
- A Deep Reinforcement Learning Strategy and Framework for Floating Waste Capture☆13Mar 13, 2025Updated last year
- ☆19Sep 16, 2025Updated 10 months ago
- [WWW 2023] Official code of "Adap-$\tau$: Adaptively Modulating Embedding Magnitude for Recommendation"☆29Jan 4, 2024Updated 2 years ago
- repository for "Exploiting Proximity-Aware Tasks for Embodied Social Navigation" paper code☆12Nov 16, 2023Updated 2 years ago
- ☆10Apr 8, 2024Updated 2 years ago
- The Software of UnionCom Algorithm☆26Jul 29, 2024Updated last year
- ☆14Jun 11, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2026] Geometric-Mean Policy Optimization☆104Jan 26, 2026Updated 5 months ago
- The repository of CLEME (EMNLP 2023) and CLEME2.0 (ACL 2025)☆12May 17, 2025Updated last year
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- source code of (quasi-)Givens Orthogonal Fine Tuning integrated to peft lib☆16Mar 13, 2025Updated last year
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- ☆25May 16, 2024Updated 2 years ago
- Code for EMNLP2023 paper "MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter".☆12Dec 27, 2023Updated 2 years ago