[ICLR'26] MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
☆54Apr 17, 2026Updated 3 months ago
Alternatives and similar repositories for MARSHAL
Users that are interested in MARSHAL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments☆25Sep 30, 2025Updated 9 months ago
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆199Mar 27, 2026Updated 3 months ago
- ☆28May 30, 2026Updated last month
- [ICLR 2026] Meta-RL Induces Exploration in Language Agents☆45Feb 1, 2026Updated 5 months ago
- Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".☆30Aug 9, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.☆22Jan 19, 2026Updated 6 months ago
- ☆120Apr 7, 2026Updated 3 months ago
- Official Implementation of HIMA (COLM'25)☆21Nov 25, 2025Updated 8 months ago
- Code for the NeurIPS 2023 Paper: Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Sta…☆29Oct 29, 2023Updated 2 years ago
- This is the official implementation of the paper "DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy".☆25Dec 19, 2025Updated 7 months ago
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- ☆28May 20, 2024Updated 2 years ago
- This repo supports integrating LLMs and communication algorithms with MARL using SMAC as the platform. It provides an end-to-end workflow…☆20Mar 8, 2025Updated last year
- A Framework for LLM-based Multi-Agent Reinforced Training and Inference☆539Apr 14, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [Paper][EMNLP 2025] RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models☆17Jan 29, 2026Updated 5 months ago
- A Collection of Competitive Text-Based Games for Language Model Evaluation and Reinforcement Learning☆411Updated this week
- ☆47Feb 8, 2024Updated 2 years ago
- Sotopia-RL: Reward Design for Social Intelligence☆52Apr 1, 2026Updated 3 months ago
- ☆14Oct 17, 2024Updated last year
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆20Dec 30, 2025Updated 6 months ago
- Official implementation for "How Should We Meta-Learn Reinforcement Learning Algorithms?"☆23Sep 7, 2025Updated 10 months ago
- ☆26Apr 21, 2026Updated 3 months ago
- ☆45Jan 9, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Hypothetical Minds is an autonomous LLM-based agent for diverse multi-agent settings, integrating a Theory of Mind module Theory of Mind …☆64Jul 13, 2024Updated 2 years ago
- A package to convert range data from ROS range topics to pointclouds☆10Jun 30, 2017Updated 9 years ago
- Offcial Repo of Paper "Eliminating Position Bias of Language Models: A Mechanistic Approach""☆23Jun 13, 2025Updated last year
- Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images☆19Jun 4, 2025Updated last year
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆95Apr 7, 2026Updated 3 months ago
- ☆11May 17, 2024Updated 2 years ago
- [NeurIPS 2025@FoRLM] R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search☆17Jan 24, 2026Updated 6 months ago
- An official implementation of "Rethinking Graph Backdoor Attacks: A Distribution-Preserving Perspective" (KDD 2024)☆12Sep 16, 2024Updated last year
- ☆70Feb 4, 2026Updated 5 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Code for EACL 26 Findings paper "I-MCTS: Enhancing Agentic AutoML via Introspective Monte Carlo Tree Search"☆13Jan 28, 2026Updated 5 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 9 months ago
- Official implementation of "Can Test-Time Scaling Improve World Foundation Model?"☆15Jul 12, 2025Updated last year
- A toolkit for automated alignment research.☆15Jul 3, 2026Updated 3 weeks ago
- [ICLR 2025 Spotlight] Weak-to-strong preference optimization: stealing reward from weak aligned model☆18Feb 24, 2025Updated last year
- [NeurIPS 2025] Official PyTorch implementation of paper "Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression".☆15Oct 24, 2025Updated 9 months ago
- [ICLR 2025] Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better☆16Feb 15, 2025Updated last year