[ICLR2026] "Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models"
☆30Feb 4, 2026Updated 5 months ago
Alternatives and similar repositories for Co-rewarding
Users that are interested in Co-rewarding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2026] "Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models"☆58Feb 4, 2026Updated 5 months ago
- [ICLR 2025] "Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond"☆16Feb 27, 2025Updated last year
- This is the PyTorch Implementation for our model VRKG4Rec (WSDM'23)☆19Mar 30, 2023Updated 3 years ago
- [arXiv:2510.06261] "AlphaApollo: A System for Deep Agentic Reasoning"☆46May 18, 2026Updated 2 months ago
- [NeurIPS 2024] "Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?"☆40Jul 18, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A System for Evaluating Reasoning Agents such as OpenClaw☆21Apr 3, 2026Updated 3 months ago
- translation of VHL repo in paddle☆25Jun 28, 2023Updated 3 years ago
- [ICLR 2025] "Noisy Test-Time Adaptation in Vision-Language Models"☆16Feb 22, 2025Updated last year
- Official implementation of "Can Test-Time Scaling Improve World Foundation Model?"☆15Jul 12, 2025Updated last year
- [NeurIPS'23] Binary Classification with Confidence Difference☆10May 13, 2024Updated 2 years ago
- [ICML 2025] "From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium"☆39Nov 23, 2025Updated 7 months ago
- [ICML 2024] Generalizing Knowledge Graph Embedding with Universal Orthogonal Parameterization☆17May 12, 2024Updated 2 years ago
- Bayesian negative sampling is the theoretically optimal negative sampling algorithm that runs in linear time.☆40Nov 7, 2025Updated 8 months ago
- [ICLR 2025] "Noisy Test-Time Adaptation in Vision-Language Models"☆13Feb 22, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [arXiv:2311.03191] "DeepInception: Hypnotize Large Language Model to Be Jailbreaker"☆176Feb 20, 2024Updated 2 years ago
- ☆18Aug 19, 2024Updated last year
- ☆19Aug 4, 2025Updated 11 months ago
- [NeurIPS 2022] "Adversarial Training with Complementary Labels: On the Benefit of Gradually Informative Attacks"☆13Nov 11, 2022Updated 3 years ago
- [ICML 2023] "On Strengthening and Defending Graph Reconstruction Attack with Markov Chain Approximation"☆22Nov 10, 2023Updated 2 years ago
- ☆43Jun 17, 2026Updated last month
- [NeurIPS2023] "Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer Learning" by Yihua Zhang*, Yimeng Zhang*,…☆14Oct 12, 2023Updated 2 years ago
- ☆15Apr 17, 2026Updated 3 months ago
- [NeurIPS 2023] Combating Bilateral Edge Noise for Robust Link Prediction☆13Nov 3, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [NeurIPS DB 2025] IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering☆46Oct 15, 2025Updated 9 months ago
- [ICASSP2025] ConcealGS: Conceal Implicit Information in 3D Gaussian Splatting☆20Jan 22, 2025Updated last year
- FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion (NeurIPS 2024 Spotlight)☆15Mar 31, 2025Updated last year
- [ICML 2025] "From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?"☆47Oct 8, 2025Updated 9 months ago
- The dataset repo of "CLCIFAR: CIFAR-Derived Benchmark Datasets with Human Annotated Complementary Labels" paper☆17May 11, 2026Updated 2 months ago
- ☆34Nov 16, 2025Updated 8 months ago
- [ICML25] Official repo for "Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond…☆24Sep 27, 2025Updated 9 months ago
- [ICML 2024] "Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models"☆59Sep 3, 2024Updated last year
- ClawNet is a governed multi-agent social network where every AI agent acts under human-granted identity, scoped authorization, and full a…☆35Jun 3, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆21Mar 17, 2025Updated last year
- [COLM 2025] SEAL: Steerable Reasoning Calibration of Large Language Models for Free☆60Apr 6, 2025Updated last year
- [ICML 2022] HousE: Knowledge Graph Embedding with Householder Parameterization☆40Feb 1, 2022Updated 4 years ago
- [ICLR 25] A novel framework for building intrinsically interpretable LLMs with human-understandable concepts to ensure safety, reliabilit…☆33Feb 5, 2026Updated 5 months ago
- ☆14Jun 5, 2024Updated 2 years ago
- source code for NeurIPS'24 paper "HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection"☆70Apr 11, 2025Updated last year
- [ICLR 2024 Spotlight] "Negative Label Guided OOD Detection with Pretrained Vision-Language Models"☆30Oct 23, 2024Updated last year