A benchmark for conversational bargaining by language models. In each 20‑round match one LLM plays buyer, one plays seller, and both hold a hidden private value. Every round they swap a short public message, then post a bid or ask; a deal clears whenever the bid meets the ask.
☆46Jun 23, 2026Updated last month
Alternatives and similar repositories for pact
Users that are interested in pact are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A multi-agent benchmark where eight LLMs play a money-driven elimination game with private transfers and a buyout endgame, and are ranked…☆19May 27, 2026Updated 2 months ago
- The BAZAAR challenges LLMs to navigate the double-auction marketplace, where buyers and sellers must make strategic decisions with incomp…☆38Jul 30, 2025Updated last year
- Systemic, uninstructed collusion among frontier LLMs in a simulated bidding environment☆19Jul 15, 2025Updated last year
- Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLM…☆90Dec 9, 2025Updated 8 months ago
- Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a sm…☆73Apr 16, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Public Goods Game (PGG) Benchmark: Contribute & Punish is a multi-agent benchmark that tests cooperative and self-interested strategies a…☆41Apr 10, 2025Updated last year
- Benchmark evaluating LLMs on their ability to create and resist disinformation. Includes comprehensive testing across major models (Claud…☆33Mar 20, 2025Updated last year
- LLM Divergent Thinking Creativity Benchmark. LLMs generate 25 unique words that start with a given letter with no connections to each oth…☆35Mar 20, 2025Updated last year
- Benchmark that evaluates LLMs using 759 NYT Connections puzzles extended with extra trick words☆237Updated this week
- Hallucinations (Confabulations) Document-Based Benchmark for RAG. Includes human-verified questions and answers.☆248Aug 7, 2025Updated last year
- Estimate the number of legal chess positions☆14Jan 14, 2021Updated 5 years ago
- This benchmark tests how well LLMs incorporate a set of 10 mandatory story elements (characters, objects, core concepts, attributes, moti…☆422Jul 20, 2026Updated 3 weeks ago
- LLM benchmark and leaderboard for narrator-bias sycophancy, opposite-narrator contradictions, and judgment consistency.☆56Aug 6, 2026Updated last week
- A multi-player tournament benchmark that tests LLMs in social reasoning, strategy, and deception. Players engage in public and private co…☆298Jan 7, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is a framework that implements various parallel reasoning strategies from the literature☆275Dec 18, 2025Updated 7 months ago
- SCREWS: A Modular Framework for Reasoning with Revisions☆27Sep 26, 2023Updated 2 years ago
- Repo for the "Exploring Messari's Crypto API" article☆10Dec 19, 2018Updated 7 years ago
- This project is the official implementation of ``Self-Supervised Graph Neural Network for Multi-Source Domain Adaptation'' in PyTorch, wh…☆12Nov 4, 2022Updated 3 years ago
- [ICLR26] Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs☆24Apr 8, 2026Updated 4 months ago
- Emotion Detection via a Deep Learning Framework☆11Feb 13, 2025Updated last year
- Code for "MSFMamba: Multi-Scale Feature Fusion State Space Model for Multi-Source Remote Sensing Image Classification"☆10Aug 26, 2024Updated last year
- Code for Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution (ACL2021)☆13Jun 2, 2021Updated 5 years ago
- SVGBench: A challenging LLM benchmark that tests knowledge, coding, physical reasoning capabilities of LLMs.☆74Feb 12, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An OpenAI API Compatible Honeypot Gateway☆27Mar 17, 2025Updated last year
- These examples demonstrate how to use the Cloudflare API within interactive Python notebooks.☆25Jun 3, 2026Updated 2 months ago
- [TMLR 2025 J2C] TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models☆54Dec 24, 2025Updated 7 months ago
- 🌸 an exploration of HTML as a medium and material☆19Updated this week
- A geometric-driven semi-supervised approach for fishing activity detection from AIS data.☆13Aug 24, 2022Updated 3 years ago
- ☆12Jan 18, 2023Updated 3 years ago
- ☆13Sep 11, 2024Updated last year
- Local-first LLM wiki — ingests your documents into a curated, citation-backed knowledge base you can read and chat with. Marimo + SQLite/…☆19Updated this week
- [ICLR 2026] A framework to "create benchmarks" and "evaluate AI co-scientists" in experimental data-driven real-world scientific research…☆18Feb 16, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆15Apr 8, 2026Updated 4 months ago
- [ARCHIVED] aurora is archived. Memory → "litectx," orchestration → "bareagent"☆18May 22, 2026Updated 2 months ago
- OpenClaw Skill: 每日论文速递 - AI/Robotics 领域自动化学术调研工具☆20Feb 24, 2026Updated 5 months ago
- A web viewer for massCode snippets☆16Sep 25, 2025Updated 10 months ago
- A general human-ai interaction platform.☆18May 27, 2026Updated 2 months ago
- Scrape every LinkedIn public profile using Scrapy (Python)☆15Jan 16, 2015Updated 11 years ago
- Loopbackd - connect to remote machines shell via Telegram/Whatsapp/SMS☆10Nov 21, 2022Updated 3 years ago