Cited 83-model x 49-benchmark LLM evaluation matrix with 18 matrix completion methods
☆40Feb 25, 2026Updated 5 months ago
Alternatives and similar repositories for llm-benchmark-matrix
Users that are interested in llm-benchmark-matrix are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Mar 10, 2026Updated 4 months ago
- Export Jupyter Notebooks to (Xe)LaTeX with Greek Support☆13Nov 25, 2018Updated 7 years ago
- ☆21Mar 11, 2025Updated last year
- ☆18Feb 22, 2025Updated last year
- A transformer that executes a one-instruction Turing-complete computer — two approaches: hand-coded weights (no training) and learned fro…☆41Mar 3, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆16Jun 17, 2025Updated last year
- ACL 2026 & NAACL 2025: Bridging Retrieval and Inference through Evidence Fusion☆14Apr 9, 2026Updated 3 months ago
- [ICLR26] AI-based scaling law discovery☆31Jan 30, 2026Updated 6 months ago
- Exploration of automated dataset selection approaches at large scales.☆55Mar 4, 2025Updated last year
- Evals meant to evaluate language models' ability to reason over long contexts.☆10Sep 12, 2024Updated last year
- 6,080-param transformer achieving 100% accuracy on 10-digit addition. Trained from scratch in 10 minutes.☆22Feb 19, 2026Updated 5 months ago
- Code for paper "ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection" (MobiSys'23)☆14Nov 1, 2023Updated 2 years ago
- GOPHI: an AMR-to-English Verbalizer☆11Feb 5, 2020Updated 6 years ago
- AlgoTune is a NeurIPS 2025 benchmark made up of 154 math, physics, and computer science problems. The goal is write code that solves each…☆112Jun 24, 2026Updated last month
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- WebSocket server implementation of OCaml☆14May 14, 2019Updated 7 years ago
- Logical inference system based on event semantics and degree semantics in formal semantics☆10Jan 22, 2023Updated 3 years ago
- Submodule of evalverse forked from [google-research/instruction_following_eval](https://github.com/google-research/google-research/tree/m…☆15May 4, 2024Updated 2 years ago
- Synthetic Data Generation for Evaluation☆16Feb 21, 2025Updated last year
- ☆31Mar 13, 2026Updated 4 months ago
- ☆15Jul 30, 2026Updated last week
- ☆10Jun 11, 2019Updated 7 years ago
- Official Code for MIMETIC^2☆13Nov 19, 2024Updated last year
- ☆19Apr 28, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- finding new ramsey bounds through scaling autoresearch☆48May 13, 2026Updated 2 months ago
- Code for ECML-PKDD 2022 Paper --- CMG: A Class-Mixed Generation Approach to Out-of-Distribution Detection☆13Oct 12, 2022Updated 3 years ago
- TPTP python library and benchmarking service☆13Jul 20, 2026Updated 2 weeks ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- Wronging a Right: Generating Better Errors to Improve Grammatical Error Detection☆16Jan 2, 2019Updated 7 years ago
- Comprehensive LLM evaluation framework: GPQA Diamond to Chatbot Arena. Tests all major models equally, easily extensible.☆17Aug 22, 2024Updated last year
- Data and all☆14Sep 30, 2019Updated 6 years ago
- [ICML 2022] Official implementation of "Score-Guided Intermediate Layer Optimization: Fast Langevin Mixing for Inverse Problems".☆12Jul 19, 2022Updated 4 years ago
- ☆18Dec 2, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [Lanser-CLI] Official Implementation of "Reinforcement Learning from Compiler and Language Server Feedback" (https://arxiv.org/abs/2510.2…☆18Jun 15, 2026Updated last month
- A curated reading list for researchers in the Philosophy of Interpretability☆17Aug 17, 2025Updated 11 months ago
- Implement the essential operators from Allens Interval Algebra, and also some metaprogramming for combinatoral operators☆13Oct 17, 2018Updated 7 years ago
- Data visualization for Inspect AI large language model evalutions.☆21Jul 15, 2026Updated 3 weeks ago
- Run Claude Code on OpenAI models☆20Jul 13, 2025Updated last year
- Automated Semantic Analysis of Discourse Markers☆11May 30, 2022Updated 4 years ago
- An abductive reasoning engine written in C++.☆13Dec 28, 2018Updated 7 years ago