☆22Jun 16, 2026Updated last month
Alternatives and similar repositories for GD2PO
Users that are interested in GD2PO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆29Jun 9, 2026Updated last month
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR☆21Apr 7, 2026Updated 3 months ago
- ☆18Feb 14, 2026Updated 5 months ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆105Jul 23, 2026Updated last week
- Official code for "Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate" (arXiv:2605.01347).☆35Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Meta-Reinforcement Learning with Self-Reflection☆33Mar 26, 2026Updated 4 months ago
- Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning☆21Jan 4, 2026Updated 7 months ago
- NeurIPS 2025 Poster☆22Oct 17, 2025Updated 9 months ago
- ☆16May 19, 2026Updated 2 months ago
- ☆16Mar 9, 2026Updated 4 months ago
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- 【ICML2026 Spotlight】 T2PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning☆53May 27, 2026Updated 2 months ago
- ☆11Feb 1, 2023Updated 3 years ago
- [CIKM 2022] Towards Automated Over-Sampling for Imbalanced Classification☆10Mar 20, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning☆18May 21, 2026Updated 2 months ago
- Multilingual Entity Linking model by BELA model☆12Jul 20, 2023Updated 3 years ago
- Code of Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed Classification☆11Nov 5, 2025Updated 8 months ago
- [IJCAI'23] Speeding Up Multi-Objective Hyperparameter Optimization by Task Similarity-Based Meta-Learning for the Tree-Structured Parzen …☆10Apr 24, 2026Updated 3 months ago
- The official repo of ICML2026 Paper: Adversarial Latent Embedding Repair for LLM Continual Learning☆19May 27, 2026Updated 2 months ago
- A pytorch implementation of our paper Image Captioning with Inherent Sentiment (ICME 2021 Oral).☆11Jul 18, 2022Updated 4 years ago
- [EMNLP 2023] ReLM: Leveraging Language Models for Enhanced Chemical Reaction Prediction.☆22Jan 28, 2024Updated 2 years ago
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated last year
- ☆11Oct 2, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official repository for "Safety in Large Reasoning Models: A Survey" - Exploring safety risks, attacks, and defenses for Large Reasoning …☆90Aug 25, 2025Updated 11 months ago
- The official repository of paper "Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models''☆113Aug 15, 2025Updated 11 months ago
- Iterating Function Systems☆11Mar 27, 2014Updated 12 years ago
- ☆13Jul 22, 2026Updated last week
- The official repo of NeurIPS2025 Paper: High-Performance Arithmetic Circuit Optimization via Differentiable Architecture Search☆19Oct 19, 2025Updated 9 months ago
- Accepted to ICLR 2025. MetaMetrics is a calibrated meta-metric designed to evaluate generation tasks across different modalities aligned …☆15Dec 30, 2024Updated last year
- Official Codebase for "Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers"☆27Jun 7, 2025Updated last year
- ☆13May 6, 2025Updated last year
- [ICML 2025] Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment (https://arxiv.org/abs/2410.02197)☆43Jun 15, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails☆22Jul 8, 2026Updated 3 weeks ago
- [NeurIPS 2022] Supervising the Multi-Fidelity Race of Hyperparameter Configurations☆14Apr 25, 2023Updated 3 years ago
- A Multi-objective Multi-fidelity acquisition function for Bayesian optimization based on EHVI method.☆14May 18, 2022Updated 4 years ago
- ☆22Mar 11, 2026Updated 4 months ago
- A user-friendly & efficient knowledge distillation framework for LLMs, supporting off-policy, on-policy (OPD), cross-tokenizer, multimoda…☆231Updated this week
- LLMPerf is a library for validating and benchmarking LLMs☆11Aug 13, 2024Updated last year
- A repository for experiments in quality-aware decoding☆18Jun 7, 2022Updated 4 years ago