Agentic RL最详细入门
☆318Aug 6, 2026Updated 3 weeks ago
Alternatives and similar repositories for Agentic-RL-Most-Detailed-Intro
Users that are interested in Agentic-RL-Most-Detailed-Intro are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codex skill for reconstructing diagram images into editable Draw.io files☆24Updated this week
- A collection of reinforcement learning notes with formula derivations, problem records, algorithm source codes and exported PDFs in Markd…☆34Jun 9, 2026Updated 2 months ago
- Local-first interview recording review reports with a Codex skill and CLI.☆79May 16, 2026Updated 3 months ago
- ☆28Jul 20, 2026Updated last month
- [ICLR 2023] This repository contains the official Pytorch implementation for the paper "Transformer-based model for symbolic regression v…☆29Aug 5, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆172Mar 18, 2026Updated 5 months ago
- Code for Tangent Model Composition for Ensembling and Continual Fine-tuning (ICCV 2023) and Tangent Transformers for Composition, Privacy…☆14May 14, 2024Updated 2 years ago
- qwen-nsa☆87Oct 14, 2025Updated 10 months ago
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,343Nov 13, 2025Updated 9 months ago
- Enhanced Search-R1 Implementation: Improved Compatibility and Modern Framework Integration☆31Dec 8, 2025Updated 8 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 10 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,304Updated this week
- Curated papers, taxonomy, benchmarks, and decision guides for credit assignment in reasoning and agentic LLM reinforcement learning.☆153Aug 3, 2026Updated 3 weeks ago
- ☆18Dec 16, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- SearchAgent-Zero: A Scalable Multi-Turn Search Agent RL Framework☆159Updated this week
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 10 months ago
- Awesome List for Agentic RL☆1,803Updated this week
- ☆81Feb 6, 2026Updated 6 months ago
- DeepSeek-V4 Lecture☆28Aug 10, 2026Updated 2 weeks ago
- The official repository of paper: MemPO: Self-Memory Policy Optimization for Long-Horizon Agents☆26Apr 10, 2026Updated 4 months ago
- Hands-on modules for LLMs: concise implementations for whiteboard-style practice.☆39May 19, 2026Updated 3 months ago
- [ICLR25] STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs☆20Jun 3, 2025Updated last year
- CS336 作业 5 实现, 附加作业里面的 dpo/rlhf 也完成了, 消融实验分析也放在飞书文档里面了, 仅供参考☆41Sep 27, 2025Updated 11 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题☆14,978Jun 14, 2026Updated 2 months ago
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards☆40Jun 1, 2026Updated 2 months ago
- 收集为大模型面试准备的手撕代码☆53Mar 15, 2026Updated 5 months ago
- Collect the awesome works evolved around reasoning models like O1/R1 in visual domain☆55Jul 21, 2025Updated last year
- ☆255Jul 17, 2026Updated last month
- XS-VID: An Extra Small Object Video Detection Dataset☆10Aug 1, 2026Updated 3 weeks ago
- This project aims to replicate mainstream open-source model architectures with limited computational resources, implementing mini models …☆292Updated this week
- Reproducing and studying RL algorithms for LLM agents, including GRPO, GSPO, DAPO, OPD, Search-R1, ReTool, ALFWorld and beyond.☆246Updated this week
- Pretrain、Posttrain、RAG、Agent等大模型相关的基础项目合集☆39Dec 7, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is a python library. Install with "python3 -m pip install rp" then run with "python3 -m rp" or just "rp". Requires python≥3.5☆13Jul 13, 2026Updated last month
- White-box Fairness Testing through Adversarial Sampling☆14Apr 16, 2021Updated 5 years ago
- Official implementation of "Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent".☆23May 23, 2025Updated last year
- Converting night into day is one of the most interesting applications in generative models, due to the great difficulty in recreating the…☆12Oct 13, 2023Updated 2 years ago
- 记录我在cs336学习时的笔记和作业☆1,090May 2, 2026Updated 3 months ago
- [ISSTA 2025] A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing☆13Feb 9, 2025Updated last year
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe☆970Aug 20, 2026Updated last week