verl: Volcano Engine Reinforcement Learning for LLMs
☆23Nov 6, 2025Updated 9 months ago
Alternatives and similar repositories for verl
Users that are interested in verl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Jan 17, 2017Updated 9 years ago
- Latent-variable Synchronous Context-Free Grammar Toolkit☆10Sep 30, 2014Updated 11 years ago
- ☆11Jun 29, 2021Updated 5 years ago
- Stack neural networks applied to hefty natural language tasks.☆15Dec 26, 2019Updated 6 years ago
- Optimizing Causal LMs through GRPO with weighted reward functions and automated hyperparameter tuning using Optuna☆60Oct 18, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Act…☆19Oct 25, 2024Updated last year
- MIRA Mini one-click local player — pip install alakazam-mira-mini; mira-mini play☆39Jul 18, 2026Updated last month
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 11 months ago
- ☆19May 25, 2026Updated 3 months ago
- Codebase for Instruction Following without Instruction Tuning☆36Sep 24, 2024Updated last year
- A simple utility to execute your deep learning scripts when there are enough idle gpus | 一个在有足够的空闲gpu时执行深度学习训练的小工具☆16Mar 22, 2022Updated 4 years ago
- Method Multi-Layer Robust Principal Component Analysis. This method was introduced in the paper Camera-Trap Images Segmentation using Mul…☆15Jan 2, 2018Updated 8 years ago
- FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments☆18Jun 4, 2026Updated 2 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICRA'25] H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps☆13Apr 10, 2025Updated last year
- ☆32Nov 14, 2025Updated 9 months ago
- Verlog: A Multi-turn RL framework for LLM agents☆73Aug 8, 2026Updated 3 weeks ago
- HTML Agent based on NexAU☆16Nov 20, 2025Updated 9 months ago
- [ICLR 2024] Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond☆23Apr 29, 2024Updated 2 years ago
- Official implementation of Browse-Master, a tool-augmented web-search agent.☆36Aug 22, 2025Updated last year
- ☆57Mar 18, 2026Updated 5 months ago
- ☆21May 4, 2026Updated 3 months ago
- AI eXplainable Inference & Search. Open Sourcing on-premise, ultra-fast latency intelligence to all.☆37Feb 28, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆17Feb 4, 2026Updated 6 months ago
- [AAAI'25 Oral] Are Expressive Models Truly Necessary for Offline RL?☆15Dec 10, 2024Updated last year
- Code for Branched Schrödinger Bridge Matching (ICLR 2026)☆17Mar 20, 2026Updated 5 months ago
- RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation☆31Apr 3, 2026Updated 4 months ago
- Code for Pushdown Layers from our EMNLP 2023 paper☆29Dec 3, 2023Updated 2 years ago
- ☆111Jul 6, 2026Updated last month
- The code and data for the paper JiuZhang3.0☆49May 26, 2024Updated 2 years ago
- Official code for paper "Surgical Post-Training: Cutting Errors, Keeping Knowledge"☆21Jun 16, 2026Updated 2 months ago
- ☆23Aug 12, 2026Updated 2 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆13Sep 24, 2024Updated last year
- ☆24Oct 10, 2025Updated 10 months ago
- [NeurIPS 2025] The official implementation of "Towards Robust Zero-Shot Reinforcement Learning"☆15Jan 2, 2026Updated 7 months ago
- A Triton-only attention backend for vLLM☆28Jul 14, 2026Updated last month
- Reverse engineered Twitter's API☆12Nov 28, 2023Updated 2 years ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆31Aug 19, 2025Updated last year
- Official Release of Multistep Quasimetric Estimation (MQE)☆20Mar 13, 2026Updated 5 months ago