qwen3-base family of models RL on gsm8k using verl, is there an RL power law on downstream tasks?
☆27Oct 19, 2025Updated 11 months ago
Alternatives and similar repositories for rl-scaling-laws
Users that are interested in rl-scaling-laws are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Oct 24, 2021Updated 4 years ago
- [ICLR 26] The official code repository for the paper "Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions".☆19Feb 9, 2026Updated 7 months ago
- Transferability of Natural Language Inference to Biomedical Question Answering☆12Mar 25, 2021Updated 5 years ago
- Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow - Tensorlfow Im…☆13Feb 2, 2019Updated 7 years ago
- Test agent capability with complex generated mazes☆16Aug 7, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆20Oct 25, 2025Updated 10 months ago
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- This repository contains implementations of the paper VUSFA☆14Mar 31, 2021Updated 5 years ago
- Improving Transformation Invariance in Contrastive Representation Learning☆12Mar 13, 2021Updated 5 years ago
- Code release for "MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning"☆11Oct 11, 2024Updated last year
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆17Feb 23, 2026Updated 6 months ago
- Repository for paper Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries☆13Jul 16, 2026Updated 2 months ago
- ☆18Mar 10, 2026Updated 6 months ago
- Code for Columbia University COMS 3997 – LLM Ethics and Foundations☆16Jan 7, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Kernel Herding for probability density estimation☆14Feb 23, 2016Updated 10 years ago
- List of papers on Self-Correction of LLMs.☆83May 19, 2026Updated 4 months ago
- All the tools that allow me to never ever open up Final Cut☆11Feb 16, 2025Updated last year
- Reinforcement Learning Finetunes Small Subnetworks in Large Language Models☆15Oct 20, 2025Updated 11 months ago
- QAlign is a new test-time alignment approach that improves language model performance by using Markov chain Monte Carlo methods.☆27Mar 2, 2026Updated 6 months ago
- PureMVC JavaScript Demo: TodoMVC☆16Oct 26, 2018Updated 7 years ago
- ☆40May 26, 2023Updated 3 years ago
- Code for paper A simple approach to case-based reasoning in knowledge bases☆34Dec 8, 2022Updated 3 years ago
- Code and Data Repo for the CoNLL Paper -- Future Lens: Anticipating Subsequent Tokens from a Single Hidden State☆21Oct 24, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Training and evaluating with OpenReward☆33Apr 28, 2026Updated 4 months ago
- Two implementations of ZeRO-1 optimizer sharding in JAX☆14Jun 11, 2023Updated 3 years ago
- HPC helper for distributed training.