The BAZAAR challenges LLMs to navigate the double-auction marketplace, where buyers and sellers must make strategic decisions with incomplete information. Each agent receives a private value and must decide how to quote based solely on the history of previous rounds. A realistic test of market intuition and strategic adaptation.
☆38Jul 30, 2025Updated last year
Alternatives and similar repositories for bazaar
Users that are interested in bazaar are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Systemic, uninstructed collusion among frontier LLMs in a simulated bidding environment☆19Jul 15, 2025Updated last year
- A multi-agent benchmark where eight LLMs play a money-driven elimination game with private transfers and a buyout endgame, and are ranked…☆19May 27, 2026Updated 3 months ago
- LLM Divergent Thinking Creativity Benchmark. LLMs generate 25 unique words that start with a given letter with no connections to each oth…☆35Mar 20, 2025Updated last year
- Documents the style side of the short-story Creative Writing LLM benchmark: we generated many short stories with a range of LLMs, then an…☆26Dec 18, 2025Updated 8 months ago
- Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a sm…☆73Apr 16, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Public Goods Game (PGG) Benchmark: Contribute & Punish is a multi-agent benchmark that tests cooperative and self-interested strategies a…☆41Apr 10, 2025Updated last year
- A benchmark for conversational bargaining by language models. In each 20‑round match one LLM plays buyer, one plays seller, and both hold…☆46Jun 23, 2026Updated 2 months ago
- Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLM…☆90Dec 9, 2025Updated 8 months ago
- LLM benchmark and leaderboard for narrator-bias sycophancy, opposite-narrator contradictions, and judgment consistency.☆59Aug 6, 2026Updated last month
- Hallucinations (Confabulations) Document-Based Benchmark for RAG. Includes human-verified questions and answers.☆249Aug 7, 2025Updated last year
- This benchmark tests how well LLMs incorporate a set of 10 mandatory story elements (characters, objects, core concepts, attributes, moti…☆436Updated this week
- Benchmark that evaluates LLMs using 759 NYT Connections puzzles extended with extra trick words☆240Updated this week
- RLVR Testing and Training☆21Aug 28, 2025Updated last year
- ☆29Aug 27, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Technical indicators☆11Feb 19, 2023Updated 3 years ago
- CEO Bench is a comprehensive evaluation framework measuring how well Large Language Models perform on executive-level decision making, st…☆24Feb 13, 2026Updated 6 months ago
- ☆14May 25, 2023Updated 3 years ago
- A browser extension to automatically extract and batch-copy transcripts from Udemy videos. Perfect for students and learners who use note…☆20Jun 14, 2026Updated 2 months ago
- gan-options-simulator☆14Apr 9, 2025Updated last year
- A cross-platform Graphical Vibe Coding Environment (VCE) for Claude Code☆34Aug 20, 2025Updated last year
- Self-contained JBIG2 compressor for PDF files☆15Jul 24, 2017Updated 9 years ago
- Property grid widget for Unity UI☆14Aug 19, 2018Updated 8 years ago
- Vigil- API security☆16Jan 8, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆17Mar 11, 2025Updated last year
- Something like objdump☆16Jan 23, 2019Updated 7 years ago
- A tool for adding function calling to llm api, available as a service by following the link☆22Aug 11, 2025Updated last year
- ☆16Feb 24, 2025Updated last year
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- Angle Project (https://code.google.com/p/angleproject/) with support for Windows Store Apps (WinRT)☆44May 2, 2014Updated 12 years ago
- LLVM Compiler Infrastructure docset for dash.☆14Jan 13, 2020Updated 6 years ago
- APEX defines how AI agents communicate with brokers, exchanges, dealers, and other execution venues. One protocol. Realtime state. Autono…☆16Jul 23, 2026Updated last month
- novoStoic2.0: Integrated Pathway Design Tool with Thermodynamic Considerations and Enzyme Selection☆11Apr 28, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- uhop☆16Dec 15, 2025Updated 8 months ago
- A powerful and user-friendly tool that generates detailed captions for your images☆21Nov 11, 2024Updated last year
- A multi-provider AI coding agent with the persona of a Tech-Priest☆18Nov 1, 2025Updated 10 months ago
- Developing K - a language model to generate OPENSCAD code from prompt☆19Dec 3, 2025Updated 9 months ago
- Interactive Social Media Simulation of Believable Human Proxies☆85Aug 28, 2026Updated last week
- Medical records you can copy and paste☆12Mar 3, 2023Updated 3 years ago
- Detecting Semantic Code Clones by Building AST-based Markov Chains Model☆10Sep 27, 2024Updated last year