☆29May 30, 2026Updated 3 months ago
Alternatives and similar repositories for spinbench
Users that are interested in spinbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated collection of papers, research blogs, open-source tools, benchmarks, and community demos for robot-use agents, including demos …☆116Updated this week
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆205Mar 27, 2026Updated 5 months ago
- MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games☆30May 10, 2026Updated 4 months ago
- A Collection of Competitive Text-Based Games for Language Model Evaluation and Reinforcement Learning☆427Aug 19, 2026Updated last month
- Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.☆24Jan 19, 2026Updated 8 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Sotopia-RL: Reward Design for Social Intelligence☆52Apr 1, 2026Updated 5 months ago
- BayesVarSel: R package to calculate Bayes factors, model choice and variable selection in linear models☆12Mar 31, 2026Updated 5 months ago
- [ICLR'26] MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs☆57Apr 17, 2026Updated 5 months ago
- ☆34Oct 31, 2024Updated last year
- Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard☆25Dec 14, 2024Updated last year
- ☆48May 10, 2026Updated 4 months ago
- [ACM MM'25] Code for the paper "Open3D-VQA: A Benchmark for Embodied Spatial Reasoning with Multimodal Large Language Model in Open Space…☆18Jul 9, 2026Updated 2 months ago
- [CVPR2024] FCS: Feature Calibration and Separation for Non-Exemplar Class Incremental Learning☆20Apr 18, 2025Updated last year
- Recycling diverse models☆47Jan 18, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆14May 9, 2024Updated 2 years ago
- Benchmarking Social Intelligence of Language Agents through Interactive Scenarios☆12Jan 4, 2025Updated last year
- ☆11Oct 11, 2023Updated 2 years ago
- ☆19Feb 9, 2026Updated 7 months ago
- Training and inference code for "Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning"☆43Jan 20, 2025Updated last year
- Code and data for the paper: Competing Large Language Models in Multi-Agent Gaming Environments☆98Jan 26, 2026Updated 7 months ago
- ☆24May 28, 2025Updated last year
- Automated Capability Discovery via Foundation Model Self-Exploration☆68Apr 16, 2026Updated 5 months ago
- Modality-Agnostic Self-Supervised Learning with Meta-Learned Masked Auto-Encoder (NeurIPS 2023)☆10Jun 5, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆112Feb 2, 2026Updated 7 months ago
- R package for working with the CCS Annotator☆13Mar 14, 2024Updated 2 years ago
- End-to-end codebase for finetuning LLMs (LLaMA 2, 3, etc.) with or without DP☆17Sep 23, 2024Updated last year
- A supervised fine-tuning method for controllable reasoning length in large language models (一种通过有监督微调实现大语言模型思考长度可控的方法)☆11May 8, 2025Updated last year
- Official Repo for MageBench: Bridging Large Multimodal Models to Agents☆21Jan 8, 2025Updated last year
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- [ICLR2026] "Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models"☆30Feb 4, 2026Updated 7 months ago
- Back-of-the-envelope stuffs in Python☆20Sep 13, 2023Updated 3 years ago
- 这是对基于大模型的多智能体系统论文的总结☆10Jun 23, 2024Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- This is for EMNLP 2024 Paper: AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction☆16Nov 4, 2024Updated last year
- Source code for the Paper "Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models"☆20Feb 1, 2026Updated 7 months ago
- ☆10Jul 4, 2024Updated 2 years ago
- ☆73Feb 4, 2026Updated 7 months ago
- "Learning Discrete and Continuous Factors of Data via Alternating Disentanglement" accepted at ICML2019☆22Aug 22, 2019Updated 7 years ago
- [NeurIPS'23] Binary Classification with Confidence Difference☆10May 13, 2024Updated 2 years ago
- How to create rational LLM-based agents? Using game-theoretic workflows!☆112Jun 8, 2025Updated last year