[ICLR'26] MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
☆56Apr 17, 2026Updated 4 months ago
Alternatives and similar repositories for MARSHAL
Users that are interested in MARSHAL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments☆25Sep 30, 2025Updated 11 months ago
- Code release for "A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play" (NeurIPS 2025), https://arxiv.or…☆64Mar 2, 2026Updated 6 months ago
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆205Mar 27, 2026Updated 5 months ago
- This repo has the code and suplementary materials of our 2024 RAL submission.☆21Nov 23, 2025Updated 9 months ago
- ☆29May 30, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code release for "Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning" (CoRL 2025), https://arxiv.o…☆24Feb 26, 2026Updated 6 months ago
- [ICLR 2026] Meta-RL Induces Exploration in Language Agents☆46Feb 1, 2026Updated 7 months ago
- Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".☆30Aug 9, 2025Updated last year
- Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.☆23Jan 19, 2026Updated 7 months ago
- ☆119Apr 7, 2026Updated 5 months ago
- Official Implementation of HIMA (COLM'25)☆22Nov 25, 2025Updated 9 months ago
- Code for the NeurIPS 2023 Paper: Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Sta…☆29Oct 29, 2023Updated 2 years ago
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- This is the official implementation of the paper "DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy".☆26Dec 19, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆28May 20, 2024Updated 2 years ago
- [ICLR 2026] A Framework for LLM-based Multi-Agent Reinforced Training and Inference☆554Aug 20, 2026Updated 2 weeks ago
- [Paper][EMNLP 2025] RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models☆18Jan 29, 2026Updated 7 months ago
- Sotopia-RL: Reward Design for Social Intelligence☆52Apr 1, 2026Updated 5 months ago
- ☆14Oct 17, 2024Updated last year
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 8 months ago
- The Rainbow Parser☆17Mar 5, 2018Updated 8 years ago
- Official implementation for "How Should We Meta-Learn Reinforcement Learning Algorithms?"☆23Sep 7, 2025Updated last year
- ☆30Apr 21, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Hypothetical Minds is an autonomous LLM-based agent for diverse multi-agent settings, integrating a Theory of Mind module Theory of Mind …☆64Jul 13, 2024Updated 2 years ago
- A package to convert range data from ROS range topics to pointclouds☆10Jun 30, 2017Updated 9 years ago
- Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images☆19Jun 4, 2025Updated last year
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆97Apr 7, 2026Updated 5 months ago
- ☆11May 17, 2024Updated 2 years ago
- [NeurIPS 2025@FoRLM] R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search☆17Jan 24, 2026Updated 7 months ago
- ☆71Feb 4, 2026Updated 7 months ago
- Code for EACL 26 Findings paper "I-MCTS: Enhancing Agentic AutoML via Introspective Monte Carlo Tree Search"☆13Jan 28, 2026Updated 7 months ago
- Official implementation of "Can Test-Time Scaling Improve World Foundation Model?"☆15Jul 12, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICLR 2025 Spotlight] Weak-to-strong preference optimization: stealing reward from weak aligned model☆18Feb 24, 2025Updated last year
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.☆10May 16, 2024Updated 2 years ago
- [ICLR 2025] Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better☆16Feb 15, 2025Updated last year
- MATE: the Multi-Agent Tracking Environment.☆48Mar 31, 2023Updated 3 years ago
- A lightweight kmer-based algorithm for designing diagnostic CRISPR assays using genome data.☆13Aug 19, 2024Updated 2 years ago
- mechanical stage☆12Aug 6, 2023Updated 3 years ago
- Codes for the paper "Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer Approach"☆15Aug 30, 2024Updated 2 years ago