[NeurIPS 25] The official implementation of SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
☆30Sep 21, 2025Updated 10 months ago
Alternatives and similar repositories for SPC
Users that are interested in SPC are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15May 27, 2025Updated last year
- Code and data release of the paper Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows☆15Oct 4, 2024Updated last year
- [IJMS 2022] DRPreter: Interpretable Anticancer Drug Response Prediction Using Knowledge-Guided Graph Neural Networks and Transformer☆15Nov 30, 2022Updated 3 years ago
- Training and inference code for "Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning"☆43Jan 20, 2025Updated last year
- Official code for paper "SPA-RL: Reinforcing LLM Agent via Stepwise Progress Attribution"☆92Sep 13, 2025Updated 11 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆19Dec 23, 2025Updated 7 months ago
- The official code of "PixelWorld: Towards Perceiving Everything as Pixels" [TMLR25]☆15Sep 12, 2025Updated 11 months ago
- [ACL2026]Code Repo for paper "Scaling Behaviors of LLM Reinforcement Learning Post-Training"☆25Jul 1, 2026Updated last month
- [ICML 2023] "Data Efficient Neural Scaling Law via Model Reusing" by Peihao Wang, Rameswar Panda, Zhangyang Wang☆14Jan 4, 2024Updated 2 years ago
- ☆28May 20, 2024Updated 2 years ago
- The official code release for Q#: Provably Optimal Distributional RL for LLM Post-Training☆20Mar 4, 2025Updated last year
- ☆30Jun 5, 2025Updated last year
- Code for the ACL 2021 paper "Structural Guidance for Transformer Language Models"☆15Sep 17, 2025Updated 10 months ago
- Self-playing Adversarial Language Game Enhances LLM Reasoning, NeurIPS 2024☆146Feb 24, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".☆30Aug 9, 2025Updated last year
- Official Repo for SwS: A Weakness-driven Problem Synthesis Framework in RL for LLM Reasoning☆42Nov 11, 2025Updated 9 months ago
- CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming☆11Nov 18, 2024Updated last year
- SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data☆24Jan 24, 2026Updated 6 months ago
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.☆10May 16, 2024Updated 2 years ago
- extract ssl certs from pcap file, only for tls-v1.2☆10Nov 3, 2020Updated 5 years ago
- [ACL 2023] Contextual Distortion Reveals Constituency: Mask Language Models are Implicit Parsers.☆14Jun 3, 2023Updated 3 years ago
- FeedbackQA: Improving Question Answering Post-Deployment with Interactive Feedback☆12Jul 13, 2022Updated 4 years ago
- ☆12Mar 22, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Benchmarking Social Intelligence of Language Agents through Interactive Scenarios☆13Jan 4, 2025Updated last year
- Code for the 2025 ACL publication "Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs"☆32Jun 25, 2025Updated last year
- Official implementation for the paper "SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physic…☆20Jul 24, 2026Updated 2 weeks ago
- [NeurIPS '25] Multi-Token Prediction Needs Registers☆32Dec 14, 2025Updated 7 months ago
- ☆47Apr 9, 2025Updated last year
- Metadata for my UK Domestic Appliance-Level Electricity (UK-DALE) dataset☆17Jul 16, 2017Updated 9 years ago
- ☆14Sep 22, 2025Updated 10 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- Implementation for Decision-focused Summarization (EMNLP2021)☆12Mar 14, 2022Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ACL 2025 Findings] Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors☆90Jun 2, 2025Updated last year
- [NeurIPS 2025] First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training☆88Oct 29, 2025Updated 9 months ago
- Suri: Multi-constraint instruction following for long-form text generation [EMNLP’24]☆27Oct 3, 2025Updated 10 months ago
- Training code of waypoint predictor in Discrete-to-Continuous VLN.☆32Mar 25, 2024Updated 2 years ago
- [CVPR 2025] LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs☆15Jun 20, 2025Updated last year
- [ICCV 2025] Boosting MLLM Reasoning with Text-Debiased Hint-GRPO☆48Jul 1, 2025Updated last year
- ☆17Apr 9, 2026Updated 4 months ago