[ICML 2026] This is the official implementation for paper HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
☆26Jun 24, 2026Updated 2 months ago
Alternatives and similar repositories for HiPER-agent
Users that are interested in HiPER-agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TailNG UI is a signal-first Angular component library with Material-like components, framework-agnostic styling, and support for Angular …☆16Updated this week
- SMC-LTL: SMC-Based LTL MultiRobot Motion Planner☆14Jul 24, 2023Updated 3 years ago
- The official implementation of the paper "A Dual-Space Framework for General Knowledge Distillation of Large Language Models".☆18Jan 4, 2026Updated 7 months ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year
- [RA-L/ICRA2024] A differentiable robot learning framework for task specifications and controller synthesis.☆17Dec 22, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Source code for the paper "Intelligent Anti-jamming based on Deep Reinforcement Learning and Transfer Learning," Siavash Barqi Janiar☆24Sep 28, 2023Updated 2 years ago
- [ICML2026] Reproduce Kimi K1.5/K2 RL algorithm and rollout system☆21Apr 9, 2026Updated 4 months ago
- A Practitioner's Guide to M(eow)ti Turn Agentic ReinfOrcement learning☆84Jan 16, 2026Updated 7 months ago
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 4 months ago
- OpenAI Gym environments for spacecraft operations problems.☆14Mar 25, 2021Updated 5 years ago
- ☆23Jun 16, 2026Updated 2 months ago
- ☆14Jun 4, 2020Updated 6 years ago
- [TSC '23 - HRL-ACRA] Implementation of our paper "Joint Admission Control and Resource Allocation of Virtual Network Embedding via Hierar…☆32May 17, 2025Updated last year
- Benchmark Test-Time Scaling of General LLM Agents☆22Apr 14, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Agentic research control plane: queue state, worker preflight, wake-gated execution, evidence sync, dashboard, alerts, and AI-generated p…☆17Jul 21, 2026Updated last month
- This is a PDOP-driven Scheduler for Optical Inter-Satellite Links enabled Global Navigation Satellite Systems.☆17Apr 25, 2025Updated last year
- On Policy Distillation Build on top of Verl☆97May 25, 2026Updated 3 months ago
- ☆18May 28, 2024Updated 2 years ago
- [IEEE TSIPN' 2022] "Scalable Perception-Action-Communication Loops with Convolutional and Graph Neural Networks", by Ting-Kuei Hu, Fernan…☆16Feb 4, 2022Updated 4 years ago
- ☆30Feb 24, 2026Updated 6 months ago
- ☆24Nov 4, 2025Updated 9 months ago
- RoboMaster☆17Jun 17, 2024Updated 2 years ago
- Real-time hallucination detection for LLMs via Geometric Drift Analysis in Hidden States.☆15Jun 3, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data☆37May 1, 2026Updated 3 months ago
- 丁立中的大模型算法工程作品集。聚焦 LLM / VLM 全链路的复现与优化,涵盖:① 预训练与微调(Pretrain / SFT / MoE / 多模态对齐);② 强化学习对齐(PPO / GRPO / DAPO,含多奖励函数与 GAE);③ 知识蒸馏(离线 KL 蒸馏、在…☆17May 26, 2026Updated 3 months ago
- ☆53Apr 9, 2025Updated last year
- Claw is a local-first AI control plane that runs powerful open models on your machine and connects to top LLM providers. It intelligently…☆22Updated this week
- AT2PO: Agentic Turn-based Policy Optimization via Tree Search☆22May 21, 2026Updated 3 months ago
- LLM budget control and cost governance for AI agents. Python library for token budgets, usage limits and guardrails for OpenAI, Anthropic…☆16Jun 17, 2026Updated 2 months ago
- Satellite dynamics Python library for orbital calculations and mission planning.☆19Sep 27, 2022Updated 3 years ago
- pyPFC: An Open-Source Python Package for Phase Field Crystal Simulations☆16Jun 25, 2026Updated 2 months ago
- A CLI tool that fetches GitHub PR diffs, analyzes them with OpenAI, and generates a Markdown code review to streamline the review process…☆11Apr 29, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This code accompanies the paper DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering.☆16Mar 20, 2023Updated 3 years ago
- A simulation tool designed to model, analyze, and optimize satellite networks. This repository provides algorithms and tools for simulati…☆18Mar 7, 2025Updated last year
- ☆18Jun 7, 2026Updated 2 months ago
- Online Adaptation of Language Models with a Memory of Amortized Contexts (NeurIPS 2024)☆78Aug 3, 2024Updated 2 years ago
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- Official code for paper "Surgical Post-Training: Cutting Errors, Keeping Knowledge"☆21Jun 16, 2026Updated 2 months ago
- Custom firmware for the Rituals the Perfume Genie 2.0☆15Jul 6, 2026Updated last month