Reproducing and studying RL algorithms for LLM agents, including GRPO, GSPO, DAPO, OPD, Search-R1, ReTool, ALFWorld and beyond.
☆464Sep 29, 2026Updated last week
Alternatives and similar repositories for agentic-rl-lab
Users that are interested in agentic-rl-lab are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- something for paper agent☆11Dec 18, 2024Updated last year
- Train deepseek r1-like reasoning LLM with ease | 轻松训练1个deepseek r1类的推理LLM☆21Feb 15, 2025Updated last year
- [CVPR 2026] SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution☆19Jul 22, 2026Updated 2 months ago
- [ICML 2025] LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models☆16Nov 4, 2025Updated 11 months ago
- 🔥This is a repository of paper list for streaming LLMs/MLLMs.☆31Apr 19, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A free tool that helps you transcribe, translate, and summarize videos in any language.☆18Feb 27, 2024Updated 2 years ago
- pretrain a wiki llm using transformers☆71Sep 1, 2024Updated 2 years ago
- ☆21Aug 16, 2023Updated 3 years ago
- [ACL 2024] Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue☆26Oct 18, 2025Updated 11 months ago
- Uncovering User Interest from Biased and Noised Watch Time in Video Recommendation. In Recsys23.☆11Jul 18, 2023Updated 3 years ago
- A Robust Provably Secure Linguistic Steganography Method with Diffusion Language Model☆17Dec 8, 2025Updated 10 months ago
- ☆27Nov 20, 2025Updated 10 months ago
- ☆24Apr 12, 2026Updated 5 months ago
- ☆15Nov 23, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [USENIX Security 2026] Membership Inference Attacks on Tokenizers of Large Language Models☆23May 22, 2026Updated 4 months ago
- A provably secure disambiguating steganography method based on grouping ambiguous pools and synchronous sampling.☆17Apr 24, 2025Updated last year
- Generative Reranker PyTerrier☆18Dec 1, 2025Updated 10 months ago
- Paper Insight - AI驱动的学术论文智能分析☆138Aug 13, 2026Updated last month
- Dataset for the paper "GenWiki: A Dataset of 1.3 Million Content-Sharing Text and Graphs for Unsupervised Graph-to-Text Generation"☆26Jan 2, 2024Updated 2 years ago
- ☆34Jul 8, 2025Updated last year
- ☆11Aug 20, 2025Updated last year
- MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos.☆34Apr 18, 2026Updated 5 months ago
- ICLR 2026 Accepted Papers Simple Analysis☆28Feb 2, 2026Updated 8 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆10Nov 1, 2021Updated 4 years ago
- [ICLR 2025] FLAT: LLM Unlearning via Loss Adjustment with Only Forget Data☆14Feb 26, 2025Updated last year
- Github Repo for Reinforced Reasoning for Embodied Planning☆19Aug 16, 2025Updated last year
- 基于文心一言和树莓派Pico的最简易桌面宠物☆87Sep 16, 2025Updated last year
- ☆12Jul 19, 2020Updated 6 years ago
- EPoG: Integrated Exploration and Sequential Manipulation on Scene Graph with LLM-based Situated Replanning☆19May 15, 2026Updated 4 months ago
- 训练一个对中文支持更好的LLaVA模型,并开源训练代码和数据。☆83Sep 6, 2024Updated 2 years ago
- ☆32Jul 13, 2026Updated 2 months ago
- ☆43May 9, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆16Mar 5, 2023Updated 3 years ago
- [NeurIPS 2025] Official Implementation of paper "LD-RoViS: Training-free robust video steganography for deterministic latent diffusion mo…☆34Mar 11, 2026Updated 6 months ago
- [Neurips 2025]StegoZip: Enhancing Linguistic Steganography Payload in Practice with Large Language Models☆33Dec 4, 2025Updated 10 months ago
- codes for Efficient Test-Time Scaling via Self-Calibration☆22Sep 13, 2025Updated last year
- 基于资金流的择时选股策略☆18May 25, 2018Updated 8 years ago
- 丁立中的大模型算法工程作品集。聚焦 LLM / VLM 全链路的复现与优化,涵盖:① 预训练与微调(Pretrain / SFT / MoE / 多模态对齐);② 强化学习对齐(PPO / GRPO / DAPO,含多奖励函数与 GAE);③ 知识蒸馏(离线 KL 蒸馏、在…☆17May 26, 2026Updated 4 months ago
- [MM'23] ProTegO: Protect Text Content against OCR Extraction Attack☆14Mar 12, 2024Updated 2 years ago