一份面向实践者的 verl 框架使用教程。verl 是字节跳动开源的大语言模型强化学习训练框架,支持 PPO、GRPO 等多种算法,以及分布式训练、AgentRL 等场景。
☆132Jul 2, 2026Updated 2 months ago
Alternatives and similar repositories for verl-tutorial
Users that are interested in verl-tutorial are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ICLR 2026 Accepted Papers Simple Analysis☆28Feb 2, 2026Updated 7 months ago
- leetcode-hot100的题目,和 Interview-code-practice-python(https://github.com/leeguandong/Interview-code-practice-python)互为一体,找工作的好帮手。☆55Oct 25, 2024Updated last year
- ☆11Dec 1, 2024Updated last year
- 个人学习的医疗大模型微调项目☆37Dec 3, 2025Updated 9 months ago
- 本项目旨在利用LangChain和大语言模型(如ZhipuAI)开发一个智能数据库问答系统。 该系统能够通过自然语言理解用户的查询请求,自动生成相应的SQL语句并执行,最后将查询结果以自然语言 形式返回用户。☆15Jul 31, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation for NeurIPS 2024 oral paper: Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation☆16Jan 27, 2025Updated last year
- Official implementation of Self-Taught Agentic Long Context Understanding (ACL 2025).☆14Sep 22, 2025Updated 11 months ago
- ☆12Feb 27, 2025Updated last year
- ☆20Oct 17, 2024Updated last year
- 智能体工作流(Agentic Workflow)应用案例的项目集合。☆17Oct 11, 2025Updated 10 months ago
- SearchAgent-Zero: A Scalable Multi-Turn Search Agent RL Framework☆162Aug 27, 2026Updated last week
- Cross-Self KV Cache Pruning for Efficient Vision-Language Inference☆10Dec 15, 2024Updated last year
- 本项目是一个基于 LangGraph和大语言模型(LLM)实现的 Agentic RAG (检索增强生成)系统。它融合了动态查询分析和自我纠错机制,能够根据用户问题的复杂度智能地选择最优的策略(直接回答、向量库检索或网络搜索),并对生成的答案进行相关性评估,从而实现更高质量…☆70Oct 21, 2025Updated 10 months ago
- Collections of RLxLM experiments using minimal codes☆14Feb 17, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 在监控画质下实现对校园自行车的重识别,包含REID模型识别,向量数据库检索,UI展示☆11Feb 13, 2024Updated 2 years ago
- [ICLR24] code for LSN☆10Oct 28, 2024Updated last year
- ML about cluster, regression, classification, and so on. As a playground. Just for fun.☆10Jun 11, 2022Updated 4 years ago
- cnn,随机森林,四种股票走势图分类☆13Feb 11, 2019Updated 7 years ago
- (ICME24) This is the offical repository of iDAT: inverse Distillation Adapter-Tuning.☆13Apr 3, 2024Updated 2 years ago
- GNNLens: A Visual Analytics Approach for Prediction Error Diagnosis of Graph Neural Networks☆11Aug 12, 2022Updated 4 years ago
- 李鲁鲁老师的 Copilot-Python 学习。和ChatGPT等大语言模型协同进化。☆10Jun 3, 2025Updated last year
- ☆10Jul 16, 2025Updated last year
- 浙江大学 学在浙大/智云课堂 辅助脚本☆57Dec 5, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Grounding Language Models for Compositional and Spatial Reasoning☆18Oct 26, 2022Updated 3 years ago
- [SIGKDD 2024] Rethinking Fair Graph Neural Networks from Re-balancing☆10Jul 15, 2024Updated 2 years ago
- This is a repo consisting of papers about LLMs' perception of their knowledge boundaries; Uncertainty Quantification; Honesty Alignment; …☆25Nov 25, 2025Updated 9 months ago
- Code for AttriBoT from "AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution"☆15Apr 21, 2025Updated last year
- Awesome List for Agentic RL☆1,834Aug 28, 2026Updated last week
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 8 months ago
- 这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。☆63Apr 13, 2025Updated last year
- The Pytorch implemetation of "FeatWalk: Enhancing Few-Shot Classification through Local View Leveraging", AAAI 2024.☆11Mar 4, 2024Updated 2 years ago
- ☆13May 1, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Agentic RL最详细入门☆329Aug 6, 2026Updated 3 weeks ago
- The implementations of paper "Reinforced Preference Optimization for Recommendation" (ReRe).☆22Nov 16, 2025Updated 9 months ago
- 复现大模型相关算法及一些学习记录☆3,510Jul 2, 2026Updated 2 months ago
- 南开大学 大数据计算及应用; NKU Big Data☆11Sep 8, 2023Updated 2 years ago
- 谦友进☆17Sep 28, 2023Updated 2 years ago
- Complete Reinforcement Learning Toolkit for Large Language Models!☆21Aug 2, 2025Updated last year
- RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction☆45Jul 15, 2026Updated last month