[ICML 2026] This is the official implementation for paper HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
☆30Jun 24, 2026Updated 3 months ago
Alternatives and similar repositories for HiPER-agent
Users that are interested in HiPER-agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TailNG UI is a signal-first Angular component library with Material-like components, framework-agnostic styling, and support for Angular …☆16Oct 2, 2026Updated last week
- Native GPU-rendered desktop development environment — pure Rust + egui + wgpu. Terminal, file editor, diff, PDF, embedded browser, LSP, a…☆20Sep 17, 2026Updated 3 weeks ago
- ☆17May 13, 2026Updated 4 months ago
- The official implementation of the paper "A Dual-Space Framework for General Knowledge Distillation of Large Language Models".☆17Jan 4, 2026Updated 9 months ago
- ☆12May 14, 2021Updated 5 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- An end-to-end security evaluation framework tailored for real-world personalized agent.☆16Feb 28, 2026Updated 7 months ago
- [ICML2026] Reproduce Kimi K1.5/K2 RL algorithm and rollout system☆21Apr 9, 2026Updated 6 months ago
- A Practitioner's Guide to M(eow)ti Turn Agentic ReinfOrcement learning☆85Jan 16, 2026Updated 8 months ago
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 5 months ago
- ☆24Jun 16, 2026Updated 3 months ago
- Benchmark Test-Time Scaling of General LLM Agents☆25Apr 14, 2026Updated 5 months ago
- ☆31May 20, 2026Updated 4 months ago
- On Policy Distillation Build on top of Verl☆101Sep 3, 2026Updated last month
- ☆18May 28, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- GadgetExplorer is a .NET command-line tool for finding potential deserialization gadget chains in managed applications. It scans one or m…☆17May 22, 2026Updated 4 months ago
- Provide reusable workflows for Unreal Engine 5.6/5.7 to simplify Blueprint, C++, UI, PCG, replication, debugging, and performance tasks.☆18Updated this week
- 丁立中的大模型算法工程作品集。聚焦 LLM / VLM 全链路的复现与优化,涵盖:① 预训练与微调(Pretrain / SFT / MoE / 多模态对齐);② 强化学习对齐(PPO / GRPO / DAPO,含多奖励函数与 GAE);③ 知识蒸馏(离线 KL 蒸馏、在…☆17May 26, 2026Updated 4 months ago
- ☆30Feb 24, 2026Updated 7 months ago
- A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data☆37May 1, 2026Updated 5 months ago
- ☆53Apr 9, 2025Updated last year
- LLM budget control and cost governance for AI agents. Python library for token budgets, usage limits and guardrails for OpenAI, Anthropic…☆18Jun 17, 2026Updated 3 months ago
- AT2PO: Agentic Turn-based Policy Optimization via Tree Search☆22May 21, 2026Updated 4 months ago
- pyPFC: An Open-Source Python Package for Phase Field Crystal Simulations☆16Sep 11, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A CLI tool that fetches GitHub PR diffs, analyzes them with OpenAI, and generates a Markdown code review to streamline the review process…☆11Apr 29, 2025Updated last year
- Label-Imbalanced and Group-Sensitive Classification under Overparameterization☆17Nov 3, 2021Updated 4 years ago
- AlphaDiana: A System for Evaluating Agentic Reasoning☆22Aug 12, 2026Updated last month
- This code accompanies the paper DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering.☆16Mar 20, 2023Updated 3 years ago
- official code repo for paper "Merging Models on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model Merging"☆26Oct 11, 2025Updated 11 months ago
- Rust parsers and editors for the Debian file formats☆20Updated this week
- Official code for paper "Surgical Post-Training: Cutting Errors, Keeping Knowledge"☆21Jun 16, 2026Updated 3 months ago
- A local dev tool where your agents are weird alien dogs. Would you let them in?☆21Jun 29, 2026Updated 3 months ago
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search in ACL'25☆106Jun 16, 2025Updated last year
- Riscrithm is a lightweight macro-assembly compiler for RISC-V, providing readable control flow, modular files, compile-time validation, m…☆28Oct 2, 2026Updated last week
- Source code to accompany research paper on training multi token prediction language models using self-distillation.☆42Feb 21, 2026Updated 7 months ago
- Official code for paper "SPA-RL: Reinforcing LLM Agent via Stepwise Progress Attribution"☆93Sep 13, 2025Updated last year
- 基于spring-boot+vue2开发的编程论坛/社区☆15Jun 6, 2026Updated 4 months ago
- 基于 DeepTutor 理念重构的个人专属 AI 导师 Agent。本项目完全抛弃了原有的执行框架,底层采用 LangChain 与 LangGraph 进行彻底重写。通过构建基于图(Graph)的认知工作流与健壮的状态管理(State Management),实现了更具…☆16Updated this week
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning☆883Feb 8, 2026Updated 8 months ago