☆30Feb 24, 2026Updated 5 months ago
Alternatives and similar repositories for OAPL
Users that are interested in OAPL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- ☆21Nov 13, 2023Updated 2 years ago
- Code release for "TempLM: Distilling Language Models into Template-Based Generators"☆14Jul 21, 2022Updated 4 years ago
- ☆19Apr 15, 2026Updated 3 months ago
- Code and data used in the paper: "Training on Incorrect Synthetic Data via RL Scales LLM Math Reasoning Eight-Fold"☆32Jun 16, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆58Mar 25, 2026Updated 4 months ago
- This is a fork of the awesome Joey-NMT with Reinforcement Learning algorithms like Policy Gradient, MRT and Advantage Actor Critic.☆27Feb 10, 2023Updated 3 years ago
- The official implementation of the paper "A Dual-Space Framework for General Knowledge Distillation of Large Language Models".☆18Jan 4, 2026Updated 7 months ago
- Official Implementation of Paper "Learning to Jump: Thinning and Thickening Latent Counts for Generative Modeling" (ICML 2023)☆10Jun 6, 2023Updated 3 years ago
- Exploring limitations of LLM-as-a-judge☆20Aug 17, 2024Updated last year
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year
- ☆25Dec 12, 2025Updated 7 months ago
- PyTorch implementation of FAIR's paper "End-to-End Memory Network", NIPS 2015☆12Oct 19, 2017Updated 8 years ago
- ☆11Oct 2, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Ludic – an LLM-RL library for the era of experience☆67Jan 9, 2026Updated 6 months ago
- ☆666Apr 7, 2026Updated 3 months ago
- A simple implementation of ReasonGenRM.☆19Apr 21, 2025Updated last year
- A repository for training nanogpt-based Chess playing language models.☆30Apr 25, 2024Updated 2 years ago
- SemiDefinite Programming Algorithm (SDPA) for Python☆12Jul 1, 2026Updated last month
- A benchmark for language models based on the UK Linguistics Olympiad☆12Mar 3, 2025Updated last year
- ☆17Jun 14, 2023Updated 3 years ago
- soft entropy estimation☆16May 29, 2026Updated 2 months ago
- Text Adventure Learning Environment Suite - Benchmark to evaluate language models on interactive text environments.☆30Jul 17, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Implementation of ``Actor-Critic Alignment for Offline-to-Online Reinforcement Learning''☆13Oct 12, 2023Updated 2 years ago
- Code for NAACL-19 paper "Relation Extraction with Temporal Reasoning Based on Memory Augmented Distant Supervision"☆10Aug 26, 2019Updated 6 years ago
- Complexity Based Prompting for Multi-Step Reasoning☆17Mar 10, 2023Updated 3 years ago
- ☆15Oct 4, 2024Updated last year
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆15Jun 28, 2025Updated last year
- Implementation of "Modeling Past and Future for Neural Machine Translation"☆15Mar 16, 2018Updated 8 years ago
- Direct preference optimization with f-divergences.☆17Nov 3, 2024Updated last year
- ☆17Apr 28, 2022Updated 4 years ago
- ☆13Nov 29, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- LockManager with deadlock detection for implementing 2PL☆13Mar 13, 2019Updated 7 years ago
- sc14 matlab application☆14Nov 24, 2014Updated 11 years ago
- This repository contains the replication of the iGSM dataset generation process from the Physics of LLM paper by Zeyuan Zhu.☆17Sep 13, 2024Updated last year
- The reproduct of the paper - Aligner: Achieving Efficient Alignment through Weak-to-Strong Correction☆21May 29, 2024Updated 2 years ago
- ☆12Oct 7, 2020Updated 5 years ago
- ☆116Jan 21, 2025Updated last year
- Implementation for Decision-focused Summarization (EMNLP2021)☆12Mar 14, 2022Updated 4 years ago