Decoupled Gradient Policy Optimization (DGPO) - Official Implementation
☆48Apr 22, 2026Updated 3 months ago
Alternatives and similar repositories for DGPO-RL
Users that are interested in DGPO-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Mass-Adaptive Soft Policy Optimization (MASPO) - Official Implementation☆57Apr 27, 2026Updated 3 months ago
- SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting☆29Jun 22, 2026Updated last month
- Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization (DynaMO) - Official Implementation☆88Apr 11, 2026Updated 3 months ago
- A toolkit for synthesizing high-quality code training data using LLM agents. It provides three independent pipelines, each producing a di…☆17Mar 10, 2026Updated 5 months ago
- A curated list of resources on on-policy distillation☆25Apr 13, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ACL 2025] Cautious Next Token Prediction☆16Jul 24, 2025Updated last year
- 【ICLR 2025 🔥】MMKE-Bench, a challenging benchmark for evaluating diverse semantic editing in real-world scenarios.☆23Apr 19, 2025Updated last year
- A user-friendly evaluation tool that encompasses all necessary components for boundary detection on PASCAL-Context and NYUD-v2 datasets.☆16Oct 31, 2023Updated 2 years ago
- (1120/1100) Labs of CS208 2023 Spring: Algorithm Design and Analysis (ADA), SUSTech. Taught by Prof. Yuhui SHI.☆11Jun 4, 2023Updated 3 years ago
- 【AAAI 2026 🔥】A benchmark that evaluates multimodel knowledge conflicts for large multimodal model☆25May 27, 2025Updated last year
- ☆22Jan 31, 2024Updated 2 years ago
- ☆31Feb 7, 2025Updated last year
- ☆13Apr 30, 2025Updated last year
- 网络流量大小预测(基于Abilene数据库)☆13Apr 10, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 一个纯 JavaScript 开源工具包,提供常用的数字、字符串、数组、对象处理和 Base64 加解密功能。☆34Mar 17, 2026Updated 4 months ago
- This is the official repo for the paper "General365: Benchmarking General Reasoning in LLMs under High Difficulty and Diversity".☆88Apr 14, 2026Updated 3 months ago
- ☆21Dec 17, 2021Updated 4 years ago
- 中译名著多译本翻译转述语料。语料仅限于用于科研教学活动。文本著作权归原著者。☆12Jul 26, 2018Updated 8 years ago
- WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario (COLING 2025)☆13Jan 5, 2025Updated last year
- ☆54May 6, 2026Updated 3 months ago
- uPU, nnPU and PN learning with Extra Trees classifier.☆20Dec 2, 2024Updated last year
- 根据维基百科历史编辑数据提取纠错语料。☆12Apr 6, 2022Updated 4 years ago
- A python script for downloading huggingface datasets and models.☆20Apr 10, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization☆82Dec 25, 2025Updated 7 months ago
- 【ICLR 2025 🔥】The code for Consistent In-Context Editing, an approach for tuning language models through contextual distributions, overco…☆56Apr 2, 2025Updated last year
- ☆51Jul 31, 2026Updated last week
- [TKDE 2024, CIKM 2022] SLA²P: Self-supervised Anomaly Detection with Adversarial Perturbation.☆39Dec 26, 2024Updated last year
- Author's PyTorch implementation of ICML'23 paper "Policy Regularization with Dataset Constraint for Offline Reinforcement Learning" for D…☆18Nov 8, 2024Updated last year
- [ICLR 2026] An official implementation of "STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence"☆43Apr 19, 2026Updated 3 months ago
- Universal Linker Editor☆46Apr 21, 2026Updated 3 months ago
- Academic Personal Homepage of Jian Tang☆15Updated this week
- A paper list for medical anomaly detection. ℱℯℯ𝓁 𝒻𝓇ℯℯ to contribute!☆59Jun 15, 2026Updated last month
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Sparse Adapter Fusion for Continual Learning in NLP - EACL 2026☆16Apr 9, 2026Updated 4 months ago
- [ EMNLP 2025 Main ] Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs☆18Nov 7, 2025Updated 9 months ago
- [ACL 2026] Dissecting Failure Dynamics in Large Language Model Reasoning☆18Apr 17, 2026Updated 3 months ago
- ☆11Apr 25, 2026Updated 3 months ago
- ☆15Mar 20, 2026Updated 4 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 9 months ago
- ☆19Feb 8, 2024Updated 2 years ago