Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper
☆19Jul 4, 2025Updated last year
Alternatives and similar repositories for answer-matching
Users that are interested in answer-matching are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆57Mar 18, 2026Updated 5 months ago
- ☆22Feb 10, 2025Updated last year
- ☆20May 23, 2025Updated last year
- Codebase from our first release.☆60Feb 17, 2026Updated 6 months ago
- ☆16Apr 26, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for the paper Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs☆69Jun 23, 2026Updated 2 months ago
- Official implementation of the ΔBelief-RL method.☆31Feb 28, 2026Updated 6 months ago
- LLM play 20questions with itself☆13Mar 31, 2023Updated 3 years ago
- Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion☆11Apr 1, 2024Updated 2 years ago
- ☆12Aug 8, 2023Updated 3 years ago
- Training vision models with full-batch gradient descent and regularization☆40Feb 14, 2023Updated 3 years ago
- A simple and efficient baseline for data attribution☆11Nov 10, 2023Updated 2 years ago
- ☆18Oct 12, 2022Updated 3 years ago
- [ECCV'24] Official Implementation of Autoregressive Visual Entity Recognizer.☆14Mar 2, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The repository for papaer "Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs"☆14Dec 16, 2024Updated last year
- ☆16Jun 19, 2026Updated 2 months ago
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025☆17Jan 12, 2026Updated 7 months ago
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆23Nov 3, 2025Updated 9 months ago
- Apertium tools☆20May 27, 2021Updated 5 years ago
- Achieve error-rate fairness between societal groups for any score-based classifier.☆19Aug 21, 2025Updated last year
- An archive of learning resources assembled by current Exun members and alumni.☆15Jun 23, 2026Updated 2 months ago
- Test-time-training on nearest neighbors for large language models☆50Apr 18, 2024Updated 2 years ago
- A reliable leaderboard algorithm for machine learning competitions☆17May 19, 2015Updated 11 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for Augment & Reduce, a scalable stochastic algorithm for large categorical distributions☆10May 16, 2018Updated 8 years ago
- Package for typesetting a book into PDF and HTML using pandoc and a bunch of other tools☆16Jul 21, 2020Updated 6 years ago
- Benchmarking of 1D pattern classification networks☆11Jul 19, 2023Updated 3 years ago
- A Kernel-Based View of Language Model Fine-Tuning https://arxiv.org/abs/2210.05643☆78Sep 4, 2023Updated 2 years ago
- Code for T-MARS data filtering☆35Aug 23, 2023Updated 3 years ago
- Pytorch ImageNet1k Loader with Bounding Boxes.☆13Jan 23, 2022Updated 4 years ago
- LLM RL envs done right, plus some training code☆17Feb 1, 2026Updated 6 months ago
- ☆31May 21, 2026Updated 3 months ago
- lecture notes written at WWU Münster | PDFs available at:☆12May 5, 2017Updated 9 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- BenchBench is a Python package to evaluate multi-task benchmarks.☆24Oct 12, 2025Updated 10 months ago
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- Code for "Exponential Family Estimation via Adversarial Dynamics Embedding" (NeurIPS 2019)☆14Nov 26, 2019Updated 6 years ago
- ☆17Feb 20, 2025Updated last year
- Data mapping framework for rust stuff☆58Mar 25, 2026Updated 5 months ago
- Algorithms for approximate attention in LLMs☆22Apr 14, 2025Updated last year