Training and evaluating with OpenReward
☆33Apr 28, 2026Updated 5 months ago
Alternatives and similar repositories for openreward-cookbook
Users that are interested in openreward-cookbook are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 5 months ago
- Rethinking the Trust Region in LLM Reinforcement Learning☆75Mar 2, 2026Updated 7 months ago
- Official implementation of the ΔBelief-RL method.☆31Feb 28, 2026Updated 7 months ago
- The best ChatGPT that $100 can buy.☆59Oct 1, 2026Updated last week
- Streamline on-policy/off-policy distillation workflows in a few lines of code☆109Aug 5, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.☆151Oct 1, 2026Updated last week
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆78Jun 29, 2026Updated 3 months ago
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Sep 23, 2026Updated 2 weeks ago
- verl: Volcano Engine Reinforcement Learning for LLMs☆23Nov 6, 2025Updated 11 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆82Sep 28, 2026Updated last week
- ☆24Jun 16, 2026Updated 3 months ago
- An approach to utomatically generating browser environment with verifiable tasks☆72Mar 24, 2026Updated 6 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 7 months ago
- Programmable chat templates for LLM training and inference.☆162Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- MIRA Mini one-click local player — pip install alakazam-mira-mini; mira-mini play☆39Jul 18, 2026Updated 2 months ago
- ☆15Jun 11, 2025Updated last year
- A Python SDK for Open Reward Standard servers and clients☆17Mar 24, 2026Updated 6 months ago
- ☆21Apr 21, 2026Updated 5 months ago
- Implementation of Continuous k-Nearest Neighbors in Python☆13Aug 11, 2023Updated 3 years ago
- gradio bbox labeling tools☆11May 12, 2023Updated 3 years ago
- Basic world models☆33Oct 30, 2025Updated 11 months ago
- The original Shared Recurrent Memory Transformer implementation☆39Aug 24, 2026Updated last month
- In-Context Reinforcement Learning for Tool Use in Large Language Models☆48Mar 26, 2026Updated 6 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Code interpreter support for o1☆31Sep 13, 2024Updated 2 years ago
- Official implementation of Continual Harness (arxiv.org/abs/2605.09998) on ARC-AGI-3☆40Jul 3, 2026Updated 3 months ago
- morgen - Model Order Reduction for Gas and Energy Networks☆19May 31, 2023Updated 3 years ago
- Modeling the volatility of commodity futures Indices☆15Mar 17, 2017Updated 9 years ago
- ☆24Sep 23, 2026Updated 2 weeks ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆42Jul 20, 2026Updated 2 months ago
- A collection of lightweight interpretability scripts to understand how LLMs think☆92Mar 18, 2026Updated 6 months ago
- ☆12Sep 14, 2026Updated 3 weeks ago
- [COLM-LLA 2026] The official implementation for paper "AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficien…☆23Aug 23, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Mixture of Experts from scratch☆14Apr 12, 2024Updated 2 years ago
- Deepseek-CoT☆10Oct 6, 2024Updated 2 years ago
- Code for Branched Schrödinger Bridge Matching (ICLR 2026)☆20Mar 20, 2026Updated 6 months ago
- ☆643May 24, 2026Updated 4 months ago
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated 4 months ago
- Batched optimisation algorithms for neural network potential driven molecular dynamics.☆17Nov 27, 2025Updated 10 months ago
- The open source implementation of the multi grouped query attention by the paper "GQA: Training Generalized Multi-Query Transformer Model…☆17Dec 11, 2023Updated 2 years ago