qwen3-base family of models RL on gsm8k using verl, is there an RL power law on downstream tasks?
☆27Oct 19, 2025Updated 9 months ago
Alternatives and similar repositories for rl-scaling-laws
Users that are interested in rl-scaling-laws are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Oct 24, 2021Updated 4 years ago
- ☆12May 23, 2024Updated 2 years ago
- [ICLR 26] The official code repository for the paper "Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions".☆18Feb 9, 2026Updated 6 months ago
- Reinforcement Learning with Pong in the Browser via TensorFlow.js☆17Jan 4, 2023Updated 3 years ago
- ☆20Oct 25, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆17Feb 23, 2026Updated 5 months ago
- [Lanser-CLI] Official Implementation of "Reinforcement Learning from Compiler and Language Server Feedback" (https://arxiv.org/abs/2510.2…☆18Jun 15, 2026Updated last month
- ☆17Mar 10, 2026Updated 5 months ago
- Code for Columbia University COMS 3997 – LLM Ethics and Foundations☆16Jan 7, 2025Updated last year
- We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their…☆22Jan 11, 2026Updated 7 months ago
- List of papers on Self-Correction of LLMs.☆82May 19, 2026Updated 2 months ago
- ☆11Nov 16, 2019Updated 6 years ago
- QAlign is a new test-time alignment approach that improves language model performance by using Markov chain Monte Carlo methods.☆27Mar 2, 2026Updated 5 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆11Jul 6, 2023Updated 3 years ago
- Code and Data Repo for the CoNLL Paper -- Future Lens: Anticipating Subsequent Tokens from a Single Hidden State☆21Oct 24, 2025Updated 9 months ago
- ☆57Aug 10, 2024Updated 2 years ago
- One command · One Microsoft login · Zero repeated auth Hours of uninterrupted access to NYU Torch from your terminal and IDE.☆15Jul 17, 2026Updated 3 weeks ago
- Training and evaluating with OpenReward☆33Apr 28, 2026Updated 3 months ago
- Random tidbits.☆13Aug 5, 2025Updated last year
- Official repository of "LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model Merging"☆31Nov 4, 2024Updated last year
- Benchmark Python and Cython code☆13Jun 13, 2014Updated 12 years ago
- Large-batch Training, Neural Network Optimization☆10Nov 8, 2019Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official Repo for SwS: A Weakness-driven Problem Synthesis Framework in RL for LLM Reasoning☆42Nov 11, 2025Updated 9 months ago
- 🔥 [ICLR 2025] Official PyTorch Model "Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark"☆27Feb 9, 2025Updated last year
- Variational autoencoder in Theano☆11Sep 14, 2017Updated 8 years ago
- A lightweight code assistant with tool-using capabilities built on HuggingFace's smolagents.☆41Jun 11, 2025Updated last year
- Official code for AAAI'20 paper "Merging Weak and Active Supervision for Semantic Parsing"☆11Dec 8, 2022Updated 3 years ago
- Prabhupadavani: A Code-mixed Speech Translation Data for 25 languages☆13Oct 12, 2022Updated 3 years ago
- Zero-Shot Open Entity Typing as Type-Compatible Grounding, EMNLP'18.☆43Jan 16, 2020Updated 6 years ago
- ☆16Aug 14, 2023Updated 2 years ago
- Optimizing Causal LMs through GRPO with weighted reward functions and automated hyperparameter tuning using Optuna☆60Oct 18, 2025Updated 9 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Fast IdEntification of State-of-The-Art models using adaptive bandit algorithms☆14Jul 15, 2022Updated 4 years ago
- This repo has been migrated to https://code.larus.se/lmas/Damerau-Levenshtein☆11Jul 21, 2023Updated 3 years ago
- The Modern Data Stack in a (Smaller) Box☆12Jan 28, 2023Updated 3 years ago
- Automatic textbook formalization of Grinberg Algebraic Combinatorics☆17Jul 28, 2026Updated last week
- FROM $f(x)$ AND $g(x)$ TO $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones☆71Jan 26, 2026Updated 6 months ago
- https://www.nlp.ecei.tohoku.ac.jp/projects/aio/☆16Aug 4, 2022Updated 4 years ago
- ☆59Aug 19, 2025Updated 11 months ago