Training and evaluating with OpenReward
☆33Apr 28, 2026Updated 2 months ago
Alternatives and similar repositories for openreward-cookbook
Users that are interested in openreward-cookbook are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fused KL divergence from hidden states for knowledge distillation☆19Apr 28, 2026Updated 2 months ago
- Official implementation of Paper "System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving"☆30Apr 17, 2026Updated 3 months ago
- Rethinking the Trust Region in LLM Reinforcement Learning☆62Mar 2, 2026Updated 4 months ago
- Official implementation of the ΔBelief-RL method.☆31Feb 28, 2026Updated 4 months ago
- The best ChatGPT that $100 can buy.☆54Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Streamline on-policy/off-policy distillation workflows in a few lines of code☆107Updated this week
- Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.☆138Apr 2, 2026Updated 3 months ago
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated 3 weeks ago
- verl: Volcano Engine Reinforcement Learning for LLMs☆22Nov 6, 2025Updated 8 months ago
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Updated this week
- Finding patch of conserved amino acid sites in 3D structure☆14Apr 13, 2025Updated last year
- ☆19Jun 16, 2026Updated last month
- An approach to utomatically generating browser environment with verifiable tasks☆63Mar 24, 2026Updated 3 months ago
- Programmable chat templates for LLM training and inference.☆133Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 4 months ago
- Using stochastic gradient descent (SGD) with explicit and implicit updates to fit large-scale statistical models.☆16Aug 21, 2014Updated 11 years ago
- Python Modeling Interface☆14Jun 18, 2026Updated last month
- MIRA Mini one-click local player — pip install alakazam-mira-mini; mira-mini play☆31Updated this week
- ☆15Jun 11, 2025Updated last year
- A JAX Research Toolkit for Visualizing, Manipulating, and Understanding Gemma Models with Multi-modal Support based on Penzai.☆95Jan 13, 2026Updated 6 months ago
- A Python SDK for Open Reward Standard servers and clients☆17Mar 24, 2026Updated 3 months ago
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning☆412May 28, 2026Updated last month
- ☆28Oct 30, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Apr 21, 2026Updated 2 months ago
- Implementation of Continuous k-Nearest Neighbors in Python☆13Aug 11, 2023Updated 2 years ago
- Basic world models☆32Oct 30, 2025Updated 8 months ago
- The original Shared Recurrent Memory Transformer implementation☆36Jul 11, 2025Updated last year
- In-Context Reinforcement Learning for Tool Use in Large Language Models☆48Mar 26, 2026Updated 3 months ago
- Code interpreter support for o1☆31Sep 13, 2024Updated last year
- Official implementation of Continual Harness (arxiv.org/abs/2605.09998) on ARC-AGI-3☆33Jul 3, 2026Updated 2 weeks ago
- [ICML 2026 GenBio Workshop] Official Implementation for "Harmonic Torsional Diffusion for Protein-Ligand Flexible Docking"☆15Jun 30, 2026Updated 2 weeks ago
- ☆37Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- morgen - Model Order Reduction for Gas and Energy Networks☆19May 31, 2023Updated 3 years ago
- [ACL 2026 Oral] From Word to World: Can Large Language Models be Implicit Text-based World Models?☆66Apr 13, 2026Updated 3 months ago
- ☆24May 26, 2026Updated last month
- Modeling the volatility of commodity futures Indices☆15Mar 17, 2017Updated 9 years ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆38Jun 23, 2026Updated 3 weeks ago
- Generic building-block toolbox for training neural networks with adaptive and recursive execution. It provides reusable components to con…☆27Jun 29, 2026Updated 3 weeks ago
- ☆18May 16, 2026Updated 2 months ago