Agentic RL最详细入门
☆372Aug 6, 2026Updated 2 months ago
Alternatives and similar repositories for Agentic-RL-Most-Detailed-Intro
Users that are interested in Agentic-RL-Most-Detailed-Intro are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A collection of reinforcement learning notes with formula derivations, problem records, algorithm source codes and exported PDFs in Markd…☆34Jun 9, 2026Updated 4 months ago
- Local-first interview recording review reports with a Codex skill and CLI.☆81May 16, 2026Updated 4 months ago
- [CVPR 2025] LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs☆15Jun 20, 2025Updated last year
- ☆29Jul 20, 2026Updated 2 months ago
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,488Nov 13, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2023] This repository contains the official Pytorch implementation for the paper "Transformer-based model for symbolic regression v…☆29Aug 5, 2026Updated 2 months ago
- ☆174Mar 18, 2026Updated 6 months ago
- qwen-nsa☆87Oct 14, 2025Updated 11 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,615Updated this week
- Enhanced Search-R1 Implementation: Improved Compatibility and Modern Framework Integration☆35Dec 8, 2025Updated 10 months ago
- ☆18Dec 16, 2025Updated 9 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 11 months ago
- ☆360Jul 22, 2026Updated 2 months ago
- SearchAgent-Zero: A Scalable Multi-Turn Search Agent RL Framework☆167Aug 27, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Awesome List for Agentic RL☆1,860Sep 15, 2026Updated 3 weeks ago
- solution for cs336-assignment1,2,5 , including colab code link and blog link.☆17Feb 20, 2026Updated 7 months ago
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 11 months ago
- 本项目旨在系统性学习和记录大语言模型(LLM)系统领域的核心知识,重点关注分布式训练、推理优化、强化学习(RLHF)对齐的原理、主流框架和工程实践。☆15Apr 9, 2026Updated 6 months ago
- ☆12Mar 19, 2024Updated 2 years ago
- ☆82Feb 6, 2026Updated 8 months ago
- 一份面向实践者的 verl 框架使用教程。verl 是字节跳动开源的大语言模型强化学习训练框架,支持 PPO、GRPO 等多种算法,以及分布式训练、AgentRL 等场景。☆149Jul 2, 2026Updated 3 months ago
- ☆56Nov 22, 2025Updated 10 months ago
- 复现大模型相关算法及一些学习记录☆3,547Jul 2, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,805Updated this week
- 黑马旅游网项目☆14Feb 19, 2023Updated 3 years ago
- 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题☆15,211Jun 14, 2026Updated 3 months ago
- DeepSeek-V4 Lecture☆31Aug 10, 2026Updated last month
- 2024CCF国际AIOps挑战赛-赛道二(GLM4):基于检索增强的运维知识问答挑战赛解决方案分享。☆14Jul 5, 2024Updated 2 years ago
- The official repository of paper: MemPO: Self-Memory Policy Optimization for Long-Horizon Agents☆29Apr 10, 2026Updated 5 months ago
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards☆40Jun 1, 2026Updated 4 months ago
- CS336 作业 5 实现, 附加作业里面的 dpo/rlhf 也完成了, 消融实验分析也放在飞书文档里面了, 仅供参考☆41Sep 27, 2025Updated last year
- 收集为大模型面试准备的手撕代码☆56Mar 15, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆32Jul 2, 2026Updated 3 months ago
- 软件学院IT项目管理临时抱佛脚的复习笔记☆10Dec 23, 2023Updated 2 years ago
- 《Smol 训练手册》:打造世界级大模型的秘诀☆60May 15, 2026Updated 4 months ago
- Source code and documents for course BUAA Software Engineering Embedded 2021. 北航计算机学院2021春季学期嵌入式软件工程代码与文档.☆30Jun 2, 2022Updated 4 years ago
- ☆290Jul 17, 2026Updated 2 months ago
- This project aims to replicate mainstream open-source model architectures with limited computational resources, implementing mini models …☆294Aug 24, 2026Updated last month
- [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)☆1,130Sep 12, 2026Updated 3 weeks ago