A curated reading list of large-language-model RL papers, organized by four research directions: Reasoning RL, Agentic RL, OPD (Off-Policy / On-Policy Distillation / Drift), and Multi-Agent*
☆25Jul 31, 2026Updated this week
Alternatives and similar repositories for awesome-agentic
Users that are interested in awesome-agentic are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- offical implementation of Jailbreak-R1☆15Jul 16, 2025Updated last year
- Curated papers, taxonomy, benchmarks, and decision guides for credit assignment in reasoning and agentic LLM reinforcement learning.☆125Updated this week
- ☆18Apr 18, 2025Updated last year
- pre-training llama3 using chinese☆13May 1, 2024Updated 2 years ago
- SAM4SS: Tailoring SAM and SAM2 for Semantic Segmentation☆11Jul 31, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Evaluation repository of wikipedia index with Dria☆10Mar 14, 2024Updated 2 years ago
- 🌟 A curated list of papers, methods, and resources on long-horizon credit assignment for agentic RL.☆51Updated this week
- ☆12Mar 24, 2023Updated 3 years ago
- FECFusion: Infrared and visible image fusion network based on fast edge convolution☆14Jan 6, 2024Updated 2 years ago
- Confidence Regulation Neurons in Language Models (NeurIPS 2024)☆16Feb 1, 2025Updated last year
- [CVPR 2023] Better “CMOS” Produces Clearer Images: Learning Space-Variant Blur Estimation for Blind Image Super-Resolution☆11Mar 19, 2024Updated 2 years ago
- 目标:构建一个更符合语言学的小而美的 llama 分词器,支持中英日三国语言☆19Jun 2, 2024Updated 2 years ago
- MEAformer: An all-MLP Transformer with Temporal External Attention for Long-term Time Series Forecasting☆13Apr 27, 2024Updated 2 years ago
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆11Apr 3, 2024Updated 2 years ago
- ☆15Jun 6, 2023Updated 3 years ago
- ☆13Oct 21, 2021Updated 4 years ago
- Automatic Jailbreaking of the Text-to-Image Generative AI Systems☆15Jun 23, 2024Updated 2 years ago
- Cross-modal Clustering with Deep Correlated Information Bottleneck Method☆11Aug 7, 2022Updated 3 years ago
- An evaluation framework for mitigating DNN backdoor attacks using data augmentations☆11Dec 10, 2020Updated 5 years ago
- Verify CPU circuits in Logisim or Verilog against MARS simulation☆10Dec 31, 2020Updated 5 years ago
- Code and full version of the paper "Hijacking Attacks against Neural Network by Analyzing Training Data"☆14Feb 28, 2024Updated 2 years ago
- Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF☆25Oct 8, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [IJCAI'24] Official code for our paper "Make Graph Neural Networks Great Again: A Generic Integration Paradigm of Topology-Free Patterns …☆15Jul 3, 2025Updated last year
- ☆22Dec 14, 2023Updated 2 years ago
- Universal DDPM for Data Synthesis☆15Jan 24, 2025Updated last year
- Automatically setup the AISHELL-4 and MSDWild dataset for usage with pyannote-database (and pyannote-audio)☆14Oct 22, 2025Updated 9 months ago
- 2022华科网安可信计算实验☆12Jun 25, 2022Updated 4 years ago
- Source code for Learn Locally Correct Globally☆14Nov 15, 2021Updated 4 years ago
- An Implementation of Deep Exhaustive Model for Nested NER☆15Jul 19, 2019Updated 7 years ago
- [CVPR 2024] "Data Poisoning based Backdoor Attacks to Contrastive Learning": official code implementation.☆16Feb 10, 2025Updated last year
- ☆42Jun 17, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- An automated pipeline that leverages LLM's meta-learning capability to iteratively design and refine red-teaming systems without human in…☆30May 24, 2026Updated 2 months ago
- Agent Learning via Early Experience - Bootstrap agent training without reward signals using DSPy☆22Mar 25, 2026Updated 4 months ago
- Code for paper <Evolving Multi-Scale Normalization for Time Series Forecasting Under Distribution Shifts>☆14Jan 3, 2025Updated last year
- A LaTeX class for books, reports or theses based on https://github.com/kenohori/thesis and https://github.com/Tufte-LaTeX/tufte-latex.☆16Oct 28, 2024Updated last year
- The official implementation for "ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning"☆15Updated this week
- Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency☆16Aug 6, 2025Updated 11 months ago
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook☆18Feb 17, 2026Updated 5 months ago