[ICLR'26] MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
☆56Apr 17, 2026Updated 3 months ago
Alternatives and similar repositories for MARSHAL
Users that are interested in MARSHAL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments☆25Sep 30, 2025Updated 10 months ago
- Code release for "A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play" (NeurIPS 2025), https://arxiv.or…☆63Mar 2, 2026Updated 5 months ago
- ☆28May 30, 2026Updated 2 months ago
- Code release for "Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning" (CoRL 2025), https://arxiv.o…☆24Feb 26, 2026Updated 5 months ago
- [ICLR 2026] Meta-RL Induces Exploration in Language Agents☆45Feb 1, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".☆30Aug 9, 2025Updated last year
- Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.☆23Jan 19, 2026Updated 6 months ago
- ☆120Apr 7, 2026Updated 4 months ago
- Official Implementation of HIMA (COLM'25)☆22Nov 25, 2025Updated 8 months ago
- Code for the NeurIPS 2023 Paper: Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Sta…☆29Oct 29, 2023Updated 2 years ago
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- This is the official implementation of the paper "DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy".☆26Dec 19, 2025Updated 7 months ago
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- ☆28May 20, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This repo supports integrating LLMs and communication algorithms with MARL using SMAC as the platform. It provides an end-to-end workflow…☆20Mar 8, 2025Updated last year
- A Framework for LLM-based Multi-Agent Reinforced Training and Inference☆547Apr 14, 2026Updated 4 months ago
- [Paper][EMNLP 2025] RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models☆18Jan 29, 2026Updated 6 months ago
- A Collection of Competitive Text-Based Games for Language Model Evaluation and Reinforcement Learning☆417Updated this week
- ☆47Feb 8, 2024Updated 2 years ago
- Sotopia-RL: Reward Design for Social Intelligence☆52Apr 1, 2026Updated 4 months ago
- ☆14Oct 17, 2024Updated last year
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆25Dec 30, 2025Updated 7 months ago
- The Rainbow Parser☆17Mar 5, 2018Updated 8 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official implementation for "How Should We Meta-Learn Reinforcement Learning Algorithms?"☆23Sep 7, 2025Updated 11 months ago
- ☆28Apr 21, 2026Updated 3 months ago
- ☆45Jan 9, 2024Updated 2 years ago
- Official code for Cross-Domain Policy Adaptation by Capturing Representation Mismatch (ICML 2024)☆15Aug 15, 2025Updated last year
- Hypothetical Minds is an autonomous LLM-based agent for diverse multi-agent settings, integrating a Theory of Mind module Theory of Mind …☆65Jul 13, 2024Updated 2 years ago
- Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images☆19Jun 4, 2025Updated last year
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆95Apr 7, 2026Updated 4 months ago
- ☆11May 17, 2024Updated 2 years ago
- [NeurIPS 2025@FoRLM] R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search☆17Jan 24, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆70Feb 4, 2026Updated 6 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 9 months ago
- Official implementation of "Can Test-Time Scaling Improve World Foundation Model?"☆15Jul 12, 2025Updated last year
- [ICLR 2025 Spotlight] Weak-to-strong preference optimization: stealing reward from weak aligned model☆18Feb 24, 2025Updated last year
- A Workbench for Autograding Retrieve/Generate Systems☆15Jun 30, 2025Updated last year
- MATE: the Multi-Agent Tracking Environment.☆48Mar 31, 2023Updated 3 years ago
- SWE-Bench-plus-plus☆25Feb 5, 2026Updated 6 months ago