This is the official implementation of paper "The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward"
☆20Feb 10, 2026Updated 7 months ago
Alternatives and similar repositories for DPH-RL
Users that are interested in DPH-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026 Oral] SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks official repos.☆28May 18, 2026Updated 4 months ago
- Simple Conversational Data Augmentation for Semi-supervised Abstractive Conversation Summarization☆10Mar 7, 2022Updated 4 years ago
- ☆21Feb 13, 2026Updated 7 months ago
- verl: Volcano Engine Reinforcement Learning for LLMs☆23Nov 6, 2025Updated 10 months ago
- Graph Neural Networks for Drug Efficacy Prediction☆12Sep 11, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated last year
- ☆12Jun 16, 2023Updated 3 years ago
- The official implemention of "Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration" (ICML 2026)☆26Feb 4, 2026Updated 7 months ago
- ParetoDrug☆12Sep 3, 2024Updated 2 years ago
- Released code for「Stance Detection on Social Media with Background Knowledge」in EMNLP2023.☆20Apr 23, 2024Updated 2 years ago
- [AAAI-25] Official repository of "Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object De…☆21Dec 27, 2024Updated last year
- DyRAMO: Dynamic Reliability Adjustment for Multi-objective Optimization☆17Mar 17, 2025Updated last year
- COMA: Efficient Structure-constrained Molecular Generation using Contractive and Margin losses☆18Oct 31, 2023Updated 2 years ago
- Improved Scaffold Hopping in Ligand-based Virtual Screening Using Neural Representation Learning☆21Mar 9, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆18Mar 13, 2024Updated 2 years ago
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning☆25Jun 25, 2025Updated last year
- Code repository for "RL Grokking Recipe: How RL Unlocks and Transfers New Algorithms in LLMs""☆35Oct 12, 2025Updated 11 months ago
- NeurIPS 2025: Discriminative Constrained Optimization for Reinforcing Large Reasoning Models☆53Mar 14, 2026Updated 6 months ago
- GOProteinGNN: Leveraging Protein Knowledge Graphs for Protein Representation Learning☆13May 13, 2025Updated last year
- Highly parallel molecular docking pipeline using Vina-GPU (dockerized) + AutoDock Vina CPU☆19Nov 19, 2024Updated last year
- ☆18Apr 11, 2023Updated 3 years ago
- Drug-Target Interactive Prediction Model using ChemBERTa and ProtBert☆16Oct 21, 2024Updated last year
- ☆14Apr 10, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆22Aug 28, 2022Updated 4 years ago
- [NeurIPS 2025] The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tun…☆40Feb 20, 2025Updated last year
- The official repository for "Multi-channel learning for integrating structural hierarchies into context-dependent molecular representatio…☆15Feb 25, 2025Updated last year
- Demonstrator for the effectiveness of transformer models, specifically the newly released ChemBERTa-2, in predicting physical-chemical pr…☆17Mar 1, 2026Updated 6 months ago
- Released code for「Self-Training with Pseudo-Label Scorer for Aspect Sentiment Quad Prediction」in ACL2024.☆23Feb 21, 2025Updated last year
- Pointer-generator transformer model and transformer model for the morphological inflection task. custom to the SIGMORPHON 2019 shared tas…☆26May 26, 2020Updated 6 years ago
- ☆10Jul 4, 2024Updated 2 years ago
- Generative Pre-Training from Molecules☆22Apr 22, 2023Updated 3 years ago
- 自举的 C 语言编译器☆10Jan 8, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS'23] Binary Classification with Confidence Difference☆10May 13, 2024Updated 2 years ago
- ☆17Jan 8, 2024Updated 2 years ago
- ☆11Mar 31, 2022Updated 4 years ago
- M5Stack Paper S3 e-ink Chinese book reader & traditional almanac calendar (農曆/黃曆) with 八字, weather, Cangjie input, and more☆17May 22, 2026Updated 3 months ago
- ☆12Oct 23, 2022Updated 3 years ago
- ☆15Mar 30, 2025Updated last year
- Byte-sized text games for code generation tasks on virtual environments☆20Jul 8, 2024Updated 2 years ago