Target Policy Optimization (JAX)
☆31Apr 18, 2026Updated 4 months ago
Alternatives and similar repositories for tpo
Users that are interested in tpo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- RL models to play Sokoban. The fastest recipe wins.☆31Jul 25, 2026Updated 3 weeks ago
- A minimal hackable implementation of policy gradient methods (GRPO, PPO, REINFORCE)☆17Feb 20, 2026Updated 6 months ago
- Collection of LLM completions for reasoning-gym task datasets☆31Jul 4, 2025Updated last year
- [EMNLP2022] Transformer-based Entity Typing in Knowledge Graphs☆15Nov 26, 2024Updated last year
- soft entropy estimation☆16May 29, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICML'26] Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning☆16Jun 1, 2026Updated 2 months ago
- Follow the Mean: controlling flow-matching generative models by shifting endpoint means with reference examples☆26May 21, 2026Updated 3 months ago
- This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.☆130Apr 7, 2026Updated 4 months ago
- Official implementation of Stackelberg PPO for morphology–control co-design.☆17Mar 17, 2026Updated 5 months ago
- An implementation of ESM2 in Equinox+JAX☆36Apr 20, 2026Updated 4 months ago
- ☆17Mar 10, 2026Updated 5 months ago
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 4 months ago
- VC-FB and MC-FB algorithms from "Zero-Shot Reinforcement Learning from Low Quality Data" (NeurIPS 2024)☆29Jan 14, 2025Updated last year
- ☆94Jun 8, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- Gym wrapper for pysc2☆10Sep 16, 2022Updated 3 years ago
- ☆15May 19, 2024Updated 2 years ago
- Evaluating language models on word puzzle games☆10Oct 25, 2024Updated last year
- Implementation for ReFactor GNNs☆15Jun 10, 2025Updated last year
- [ACL2025 Best Paper] Language Models Resist Alignment☆52Jun 11, 2025Updated last year
- [CVPR'26] SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization☆22Feb 19, 2026Updated 6 months ago
- Symbol-Equivariant Recurrent Reasoning Model☆16Mar 4, 2026Updated 5 months ago
- This repositories contains the reference implementation for the Sparse Delta Memory paper.More precisely, it contains the model definitio…☆36Jul 9, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICML2026] Official JAX code for Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying☆16Jul 3, 2026Updated last month
- An open-source Rust implementation of Shape Complementarity (SC)☆20Oct 2, 2025Updated 10 months ago
- Convertible environment for reinforcement learning with Kerbal Space Program☆11Dec 8, 2022Updated 3 years ago
- This repository is a package to provide SAIT Machine Learning Force Field(MLFF) Framework☆40Oct 25, 2023Updated 2 years ago
- Framework for modified sampling from biomolecular generative models☆25Updated this week
- World-Gymnast: Training Robots with Reinforcement Learning in a World Model☆48Feb 11, 2026Updated 6 months ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆25Jun 26, 2026Updated last month
- ☆58May 26, 2026Updated 2 months ago
- ☆10Jul 14, 2018Updated 8 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Material for tutorial "Hybrid techniques for knowledge-based NLP: Knowledge graphs meet machine learning and all their friends" at KCAP…☆16Dec 4, 2017Updated 8 years ago
- Archives for Triton Inference Server Practices☆15Feb 28, 2022Updated 4 years ago
- ☆17Mar 10, 2026Updated 5 months ago
- Recommendation models that use binary rather than floating point operations at prediction time.☆21Sep 18, 2017Updated 8 years ago
- experiment☆12Jan 1, 2023Updated 3 years ago
- ☆13Sep 24, 2024Updated last year
- Code for Explainable Synthesizability Prediction of Inorganic Crystal Polymorphs using Large Language Models☆13Feb 19, 2025Updated last year