☆30Feb 24, 2026Updated 7 months ago
Alternatives and similar repositories for OAPL
Users that are interested in OAPL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- ☆31May 20, 2026Updated 4 months ago
- Code release for "TempLM: Distilling Language Models into Template-Based Generators"☆14Jul 21, 2022Updated 4 years ago
- ☆20Apr 15, 2026Updated 5 months ago
- Data and code for the preprint "In-Context Learning with Long-Context Models: An In-Depth Exploration"☆44Aug 20, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code and data used in the paper: "Training on Incorrect Synthetic Data via RL Scales LLM Math Reasoning Eight-Fold"☆32Jun 16, 2024Updated 2 years ago
- ☆58Mar 25, 2026Updated 6 months ago
- This is a fork of the awesome Joey-NMT with Reinforcement Learning algorithms like Policy Gradient, MRT and Advantage Actor Critic.☆27Feb 10, 2023Updated 3 years ago
- The official implementation of the paper "A Dual-Space Framework for General Knowledge Distillation of Large Language Models".☆17Jan 4, 2026Updated 9 months ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year
- ☆26Dec 12, 2025Updated 9 months ago
- [ICML2026] Reproduce Kimi K1.5/K2 RL algorithm and rollout system☆21Apr 9, 2026Updated 5 months ago
- Safe SLAC, an algorithm for safe cost-constrained reinforcement learning in high-dimensional POMDPs.☆13Mar 1, 2023Updated 3 years ago
- PyTorch implementation of FAIR's paper "End-to-End Memory Network", NIPS 2015☆12Oct 19, 2017Updated 8 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆11Oct 2, 2023Updated 3 years ago
- ☆15Oct 5, 2023Updated 3 years ago
- ☆714Apr 7, 2026Updated 5 months ago
- A repository for training nanogpt-based Chess playing language models.☆30Apr 25, 2024Updated 2 years ago
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 5 months ago
- Towards Formalizing RL Theory☆58Aug 21, 2026Updated last month
- ☆70Aug 28, 2026Updated last month
- A benchmark for language models based on the UK Linguistics Olympiad☆12Mar 3, 2025Updated last year
- The repository for the paper "Predicting in-hospital mortality by combining clinical notes with time-series data"☆12May 23, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆16May 22, 2025Updated last year
- ☆18Jun 14, 2023Updated 3 years ago
- ☆18Mar 3, 2025Updated last year
- soft entropy estimation☆16May 29, 2026Updated 4 months ago
- This is the unofficial implementation of LEMON (ICLR'2024).☆13Apr 14, 2024Updated 2 years ago
- Implementation of ``Actor-Critic Alignment for Offline-to-Online Reinforcement Learning''☆13Oct 12, 2023Updated 2 years ago
- Code for NAACL-19 paper "Relation Extraction with Temporal Reasoning Based on Memory Augmented Distant Supervision"☆10Aug 26, 2019Updated 7 years ago
- Text Adventure Learning Environment Suite - Benchmark to evaluate language models on interactive text environments.☆32Sep 10, 2026Updated 3 weeks ago
- Official code for ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning (AAAI'24)☆16Feb 10, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Complexity Based Prompting for Multi-Step Reasoning☆17Mar 10, 2023Updated 3 years ago
- ☆15Oct 4, 2024Updated 2 years ago
- Code space for L4DC paper "State-wise Safe Reinforcement Learning With Pixel Observations"☆11Apr 5, 2024Updated 2 years ago
- Work in progress save editor for Monster Hunter: World☆11Aug 15, 2018Updated 8 years ago
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆14Jun 28, 2025Updated last year
- Implementation of "Modeling Past and Future for Neural Machine Translation"☆15Mar 16, 2018Updated 8 years ago
- ☆17Apr 28, 2022Updated 4 years ago