Training and evaluating with OpenReward
☆33Apr 28, 2026Updated 4 months ago
Alternatives and similar repositories for openreward-cookbook
Users that are interested in openreward-cookbook are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 4 months ago
- Official implementation of Paper "System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving"☆30Apr 17, 2026Updated 5 months ago
- Official implementation of the ΔBelief-RL method.☆31Feb 28, 2026Updated 6 months ago
- The best ChatGPT that $100 can buy.☆58Updated this week
- Streamline on-policy/off-policy distillation workflows in a few lines of code☆109Aug 5, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.☆147Jul 22, 2026Updated last month
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated 2 months ago
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Updated this week
- verl: Volcano Engine Reinforcement Learning for LLMs☆23Nov 6, 2025Updated 10 months ago
- ☆24Jun 16, 2026Updated 3 months ago
- An approach to utomatically generating browser environment with verifiable tasks☆69Mar 24, 2026Updated 5 months ago
- Programmable chat templates for LLM training and inference.☆156Updated this week
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 6 months ago
- Using stochastic gradient descent (SGD) with explicit and implicit updates to fit large-scale statistical models.☆16Aug 21, 2014Updated 12 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Jun 11, 2025Updated last year
- A Python SDK for Open Reward Standard servers and clients☆17Mar 24, 2026Updated 5 months ago
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning☆456May 28, 2026Updated 3 months ago
- ☆21Apr 21, 2026Updated 4 months ago
- ☆27Oct 30, 2025Updated 10 months ago
- Implementation of Continuous k-Nearest Neighbors in Python☆13Aug 11, 2023Updated 3 years ago
- Basic world models☆33Oct 30, 2025Updated 10 months ago
- The original Shared Recurrent Memory Transformer implementation☆39Aug 24, 2026Updated 3 weeks ago
- In-Context Reinforcement Learning for Tool Use in Large Language Models☆48Mar 26, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code interpreter support for o1☆31Sep 13, 2024Updated 2 years ago
- Model-Based Visual Planning with Self-Supervised Functional Distances (ICLR 2021)☆20Jul 31, 2021Updated 5 years ago
- morgen - Model Order Reduction for Gas and Energy Networks☆19May 31, 2023Updated 3 years ago
- [ACL 2026 Oral] From Word to World: Can Large Language Models be Implicit Text-based World Models?☆74Apr 13, 2026Updated 5 months ago
- ☆13Mar 9, 2024Updated 2 years ago
- Modeling the volatility of commodity futures Indices☆15Mar 17, 2017Updated 9 years ago
- ☆20Nov 13, 2022Updated 3 years ago
- A collection of lightweight interpretability scripts to understand how LLMs think☆92Mar 18, 2026Updated 6 months ago
- [COLM-LLA 2026] The official implementation for paper "AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficien…☆24Aug 23, 2026Updated 3 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Deepseek-CoT☆10Oct 6, 2024Updated last year
- Code for Branched Schrödinger Bridge Matching (ICLR 2026)☆20Mar 20, 2026Updated 5 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆52May 30, 2026Updated 3 months ago
- ☆637May 24, 2026Updated 3 months ago
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated 3 months ago
- Official Codebase: LT2: Linear-Time Looped Transformers.☆62Jul 27, 2026Updated last month
- Official implementation of "Data Mixture Inference: What do BPE tokenizers reveal about their training data?"☆23May 15, 2025Updated last year