☆211Dec 20, 2024Updated last year
Alternatives and similar repositories for SimpleBench
Users that are interested in SimpleBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Jul 5, 2024Updated 2 years ago
- Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLM…☆89Dec 9, 2025Updated 7 months ago
- ☆136May 2, 2025Updated last year
- LiveBench: A Challenging, Contamination-Free LLM Benchmark☆1,262Updated this week
- utilities for batched llm calls with retries☆51Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a sm…☆72Apr 16, 2026Updated 3 months ago
- Public repository for the Remote Labor Index (RLI)☆75Nov 3, 2025Updated 8 months ago
- Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and …☆48Jul 19, 2026Updated last week
- Common Voice Generator using Speech Synthesizer☆14Jul 28, 2021Updated 4 years ago
- This project benchmarks 41 open-source large language models across 19 evaluation tasks using the lm-evaluation-harness library.☆102Sep 5, 2025Updated 10 months ago
- Evolutionary Search for expert-level performance on any task with environmental feedback☆14Oct 12, 2025Updated 9 months ago
- world's stupidest moe llm in 103M parameters☆20Jul 18, 2025Updated last year
- Coding problems used in aider's polyglot benchmark☆222Dec 22, 2024Updated last year
- Aidan Bench attempts to measure <big_model_smell> in LLMs.☆320Jun 26, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Sandboxed tools and JS runtime for AI agents☆17Jul 13, 2026Updated last week
- Knowledge Graph-based Retrieval-Augmented Generation for Schema Matching☆18Jun 17, 2026Updated last month
- Train your own SOTA deductive reasoning model☆111Mar 6, 2025Updated last year
- Analysis code for Neurips 2025 paper "SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks"☆56Aug 6, 2025Updated 11 months ago
- ☆50Jul 4, 2025Updated last year
- Clue inspired puzzles for testing LLM deduction abilities☆47Mar 19, 2026Updated 4 months ago
- PyTorch code for System-1.x: Learning to Balance Fast and Slow Planning with Language Models☆25Jul 22, 2024Updated 2 years ago
- BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution☆61Oct 13, 2025Updated 9 months ago
- Exploring the Limitations of Large Language Models on Multi-Hop Queries☆33Mar 2, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Drift detection module for machine learning pipelines.☆25Jun 21, 2023Updated 3 years ago
- A benchmark for emotional intelligence in large language models☆444Jul 26, 2024Updated 2 years ago
- This repository is an implementation of inferring the PaliGemma Vision Language Model on Android using Hugging Face-Gradio Client API for…☆20May 16, 2026Updated 2 months ago
- ☆121Jun 24, 2026Updated last month
- ☆17Dec 5, 2023Updated 2 years ago
- ☆51Oct 1, 2025Updated 9 months ago
- Anthropic's Contextual Retrieval implementation with visual chunk comparison. Preview context enrichment before/after embedding.☆30Sep 25, 2025Updated 10 months ago
- ☆84Mar 11, 2025Updated last year
- prinzbench is a private benchmark that ranks LLMs based on their ability to conduct legal research and analysis and locate obscure public…☆123Jul 18, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- slowly building a set of infinite riddle generators for data-hungry methods☆14Nov 15, 2022Updated 3 years ago
- Public repository containing METR's DVC pipeline for eval data analysis☆303Mar 6, 2026Updated 4 months ago
- Fast Image Integrity Checker: Scan for corrupted images using Nvidia DALI☆22Jun 20, 2021Updated 5 years ago
- Shared Lurk source code, including tests and library code.☆18Mar 3, 2024Updated 2 years ago
- ☆37Mar 2, 2026Updated 4 months ago
- Bayesian Visual Working Memory in Python.☆13Mar 28, 2020Updated 6 years ago
- A framework for pitting LLMs against each other in an evolving library of games ⚔☆35Apr 17, 2025Updated last year