☆28Oct 30, 2025Updated 9 months ago
Alternatives and similar repositories for onpolicydistillation
Users that are interested in onpolicydistillation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An RL environment similar to Cognition's SWE-Grep☆17Mar 10, 2026Updated 5 months ago
- ☆47Oct 23, 2025Updated 9 months ago
- Code for the multi-agent computer use project.☆21Jul 3, 2026Updated last month
- Second Generation of Large Language Models☆21Jun 30, 2025Updated last year
- Implementation of Direct Preference Optimization☆17Jul 17, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆15Jun 19, 2025Updated last year
- Ludic – an LLM-RL library for the era of experience☆68Aug 9, 2026Updated last week
- ☆11Feb 9, 2024Updated 2 years ago
- ☆24Mar 23, 2026Updated 4 months ago
- ☆15Jun 26, 2026Updated last month
- ☆10Jul 15, 2024Updated 2 years ago
- Training and evaluating with OpenReward☆33Apr 28, 2026Updated 3 months ago
- Understand what physics/algorithms do transformers learn internally when trained on planetary motion☆48Feb 9, 2026Updated 6 months ago
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆36Mar 20, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Simple demo ilustrating the use of LSTM neural network to predict daily changes in the Ethereum cryptocurrency☆10Jan 23, 2018Updated 8 years ago
- ☆10Nov 6, 2024Updated last year
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions☆16Jun 28, 2026Updated last month
- A lightweight computational physics framework, based on the organization of turboWAVE. Implements a "Simulation, PhysicsModule, ComputeTo…☆12Jul 23, 2026Updated 3 weeks ago
- KnowMAN: Weakly Supervised Multinomial Adversarial Networks☆12Nov 9, 2021Updated 4 years ago
- Code for Negation Neglect☆16May 22, 2026Updated 2 months ago
- Karpathy's llama2.c transpiled to MLX for Apple Silicon☆14Dec 28, 2023Updated 2 years ago
- a benchmark to evaluate the situated inductive reasoning☆19Jan 7, 2025Updated last year
- Collection of LLM completions for reasoning-gym task datasets☆31Jul 4, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Adaptable Agent Populations via a Generative Model of Policies☆12Oct 14, 2021Updated 4 years ago
- ☆13Apr 16, 2025Updated last year
- ☆120Apr 7, 2026Updated 4 months ago
- An implementation of Self-Calibrating Conformal Prediction, accepted to Neurips 2024. SC-CP combines Venn-Abers calibration and conformal…☆14Apr 16, 2026Updated 4 months ago
- A cog model for the all-mpnet-base-v2 sentence-transformers embedding model.☆15Jan 3, 2024Updated 2 years ago
- A public repo that contains integrations for Argilla and LlamaIndex.☆17Oct 10, 2024Updated last year
- Expand -> Retrieve -> Rerank - simple method with strong results on BRIGHT benchmark☆22Aug 22, 2025Updated 11 months ago
- distilled Self-Critique refines the outputs of a LLM with only synthetic data☆11Apr 11, 2024Updated 2 years ago
- Audit any agent decision across its past, present, and future, on one typed graph.☆21Jul 20, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Because it's there.☆16Sep 22, 2024Updated last year
- QLoRA for Masked Language Modeling☆23Sep 11, 2023Updated 2 years ago
- Multilingual acoustic word embedding approaches applied and evaluated on GlobalPhone data.☆11Nov 3, 2020Updated 5 years ago
- ☆25Oct 10, 2025Updated 10 months ago
- ☆12Mar 3, 2023Updated 3 years ago
- This is MPE-pytorch, fix some bugs.☆11Apr 26, 2020Updated 6 years ago
- Einsum-like high-level array sharding API for JAX☆35Jul 16, 2024Updated 2 years ago