A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the adoption of best practices in LLM assessment, and critically assess the effectiveness of these evaluation methods.
☆196Jul 6, 2026Updated 3 weeks ago
Alternatives and similar repositories for LLMEvaluation
Users that are interested in LLMEvaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A framework for few-shot evaluation of autoregressive language models.☆13Feb 14, 2024Updated 2 years ago
- First token cutoff sampling inference example☆30Jan 15, 2024Updated 2 years ago
- Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard a…☆2,130Dec 3, 2025Updated 7 months ago
- [ACL'25] Official Code for LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs☆317Jul 13, 2025Updated last year
- Python client to integrate Cleanlab Codex with your AI Agent☆19Nov 19, 2025Updated 8 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is code for How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis☆18Nov 5, 2025Updated 8 months ago
- ☆24Dec 12, 2024Updated last year
- Python client library for Cleanlab Trustworthy Language Model☆24Dec 9, 2025Updated 7 months ago
- Table detection with Florence.☆15Jul 11, 2024Updated 2 years ago
- Google 공식 Rouge Implementation을 한국어에서 사용할 수 있도록 처리☆17Jan 3, 2024Updated 2 years ago
- Text Match Cut Video Generator Web App☆37Feb 19, 2026Updated 5 months ago
- ☆10Oct 28, 2024Updated last year
- Sculpt: Structuring unstructured data with LLMs☆40Sep 22, 2025Updated 10 months ago
- ☆13May 30, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆80Jun 5, 2024Updated 2 years ago
- Tensorflow object detection api in single line☆12Dec 23, 2021Updated 4 years ago
- ↔️ T5 Machine Translation from English to Korean☆18Aug 11, 2022Updated 3 years ago
- An analysis of the systemic impact, on the Web, of AI systems, and in particular ones based on Machine Learning models, and the role that…☆30Sep 17, 2025Updated 10 months ago
- Evaluation tools shared across anserini, pyserini, and pygaggle☆36Jul 14, 2026Updated 2 weeks ago
- Experiments to assess SPADE on different LLM pipelines.☆17Apr 7, 2024Updated 2 years ago
- The goal of this repository is to accelerate Azure OpenAI service adoption and put an enterprise governance structure around it using Azu…☆12Sep 13, 2023Updated 2 years ago
- A curated list of awesome resources, libraries, frameworks, and tools for multi-agent systems (MAS) research and development.☆33Feb 17, 2025Updated last year
- Lightweight ngram random text generator☆12Jul 11, 2014Updated 12 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Linear Relational Embeddings (LREs) and Linear Relational Concepts (LRCs) for LLMs in PyTorch☆11Aug 7, 2024Updated last year
- Academic Evaluation Modules☆11Updated this week
- EmbedRank implemented in Python.☆15Jun 17, 2024Updated 2 years ago
- [ACL 2024 Findings] Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruning☆13Sep 2, 2024Updated last year
- R code for cost-effectiveness model evaluating directly acting oral anticoagulants (DOACs) for prevention of stroke in atrial fibrillatio…☆12Jan 14, 2020Updated 6 years ago
- mcp wrapper for openai built-in tools☆12Mar 13, 2025Updated last year
- Data and info for the paper "ParaDetox: Text Detoxification with Parallel Data"☆34Apr 2, 2025Updated last year
- SeeGULL is a broad-coverage stereotype dataset in English containing stereotypes about identity groups spanning 178 countries across 8 di…☆38Sep 25, 2023Updated 2 years ago
- A framework for few-shot evaluation of language models.☆13,443Jul 13, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Inspect: A framework for large language model evaluations☆2,424Updated this week
- The GopherCon 2021 "Production AI with Go" workshop materials.☆13Dec 6, 2021Updated 4 years ago
- Data-Driven Evaluation for LLM-Powered Applications☆515Jan 22, 2025Updated last year
- This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception o…☆29Jul 9, 2025Updated last year
- 🌊 ADHD-friendly productivity app. One next step at a time, flow-based workflows.☆16Feb 9, 2026Updated 5 months ago
- BERT Sentiment Classification on the IMDb Large Movie Review Dataset.☆17Sep 8, 2022Updated 3 years ago
- A collection of notes on Data Science☆32Jun 22, 2026Updated last month