SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
☆29Jun 22, 2026Updated 2 months ago
Alternatives and similar repositories for SCOPE
Users that are interested in SCOPE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Mass-Adaptive Soft Policy Optimization (MASPO) - Official Implementation☆57Apr 27, 2026Updated 3 months ago
- Official Implementation of Trajectory-Refined Distillation☆35Jun 9, 2026Updated 2 months ago
- A curated list of resources on on-policy distillation☆25Apr 13, 2026Updated 4 months ago
- Source code for SWIFT, an efficient reward model.☆21Jan 13, 2026Updated 7 months ago
- On Policy Distillation Build on top of Verl☆95May 25, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆23Apr 5, 2026Updated 4 months ago
- 📊 A simple command-line utility for querying and monitoring GPU status☆14Aug 3, 2023Updated 3 years ago
- ☆18Feb 2, 2022Updated 4 years ago
- ☆80May 8, 2026Updated 3 months ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- Code for the paper "Data Attribution for Text-to-Image Models by Unlearning Synthesized Images."☆17May 23, 2025Updated last year
- ☆581May 10, 2026Updated 3 months ago
- Template for project development.☆14Updated this week
- Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe☆57Jun 10, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Awesome List for On-Policy Distillation☆841Jul 31, 2026Updated 3 weeks ago
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards☆40Jun 1, 2026Updated 2 months ago
- A curated collection of papers and resources on On-Policy Distillation for Large Language Models.☆521Aug 12, 2026Updated last week
- AutoThink is a reinforcement learning framework designed to equip R1-style language models with adaptive reasoning capabilities. Instead …☆52Oct 14, 2025Updated 10 months ago
- Course project. A implementation of Graph Wavelet Neural Network (ICLR 2019)☆11Jan 6, 2020Updated 6 years ago
- ☆18Mar 16, 2026Updated 5 months ago
- A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models☆734Aug 15, 2026Updated last week
- End-to-end codebase for finetuning LLMs (LLaMA 2, 3, etc.) with or without DP☆17Sep 23, 2024Updated last year
- Official code for "Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate" (arXiv:2605.01347).☆36Jul 31, 2026Updated 3 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [NeurIPS 2025] Reasoning MLLM, Share-GRPO, advantage vanishing, sparse reward☆38Sep 19, 2025Updated 11 months ago
- ICML 2024 - Self-Driven Entropy Aggregation for Byzantine-Robust Heterogeneous Federated Learning☆10Jul 16, 2024Updated 2 years ago
- Structure From Motion☆10Nov 22, 2022Updated 3 years ago
- This repository contains the code for the paper The Open Proof Corpus: Building a Large-Scale, Human-Validated Dataset of LLM-Generated P…☆18Aug 4, 2025Updated last year
- ☆12Jul 30, 2025Updated last year
- Official repository for "Safety in Large Reasoning Models: A Survey" - Exploring safety risks, attacks, and defenses for Large Reasoning …☆90Aug 25, 2025Updated 11 months ago
- Paper: “MEMRL: SELF-EVOLVING AGENTS VIA RUNTIME REINFORCEMENT LEARNING ON EPISODIC MEMORY” Open-Source Code☆169Jul 18, 2026Updated last month
- This repository is the official implementation for VISD.☆23May 17, 2026Updated 3 months ago
- Reinforcement Learning via Self-Distillation (SDPO)☆1,073Jul 1, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official PyTorch implementation for the ICML 2023 paper "Out-of-Distribution Generalization of Federated Learning via Implicit Invariant …☆14Oct 31, 2023Updated 2 years ago
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe☆945Updated this week
- 【NeurIPS 2021 Spotlight】 "Torwards Gradient-based Bilevel Optimization with non-convex Followers and Beyond"☆11Mar 28, 2022Updated 4 years ago
- ☆22Jun 16, 2026Updated 2 months ago
- ☆73Feb 1, 2026Updated 6 months ago
- ☆23May 15, 2020Updated 6 years ago
- This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.☆18Jul 14, 2026Updated last month