A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the adoption of best practices in LLM assessment, and critically assess the effectiveness of these evaluation methods.
☆201Sep 1, 2026Updated last week
Alternatives and similar repositories for LLMEvaluation
Users that are interested in LLMEvaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A framework for few-shot evaluation of autoregressive language models.☆13Feb 14, 2024Updated 2 years ago
- First token cutoff sampling inference example☆30Jan 15, 2024Updated 2 years ago
- RACF (Recency Aware Collaborative Filtering) is the implementation of the "Recency Aware Collaborative Filtering for Next Basket Recommen…☆25Feb 15, 2025Updated last year
- Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard a…☆2,143Dec 3, 2025Updated 9 months ago
- [ACL'25] Official Code for LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs☆317Jul 13, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Python client to integrate Cleanlab Codex with your AI Agent☆19Nov 19, 2025Updated 9 months ago
- ☆24Dec 12, 2024Updated last year
- Google 공식 Rouge Implementation을 한국어에서 사용할 수 있도록 처리☆17Jan 3, 2024Updated 2 years ago
- Text Match Cut Video Generator Web App☆37Feb 19, 2026Updated 6 months ago
- ☆17Nov 23, 2023Updated 2 years ago
- ☆14May 30, 2024Updated 2 years ago
- ☆10May 22, 2019Updated 7 years ago
- ☆81Jun 5, 2024Updated 2 years ago
- It's a cooler way to store simple linear models.☆26Jul 15, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Tensorflow object detection api in single line☆12Dec 23, 2021Updated 4 years ago
- ↔️ T5 Machine Translation from English to Korean☆18Aug 11, 2022Updated 4 years ago
- My personal frontpage app☆112Aug 23, 2026Updated 2 weeks ago
- An analysis of the systemic impact, on the Web, of AI systems, and in particular ones based on Machine Learning models, and the role that…☆30Sep 17, 2025Updated 11 months ago
- An assignment for CMU CS11-711 Advanced NLP, building NLP systems from scratch☆171Dec 15, 2022Updated 3 years ago
- Deep Learning Paper Implementations in PyTorch☆17Mar 26, 2025Updated last year
- Evaluation tools shared across anserini, pyserini, and pygaggle☆36Aug 24, 2026Updated 2 weeks ago
- ☆58Apr 18, 2026Updated 4 months ago
- A curated list of awesome resources, libraries, frameworks, and tools for multi-agent systems (MAS) research and development.☆33Feb 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Multiple ways to model user preference in recommender systems☆18May 2, 2024Updated 2 years ago
- Linear Relational Embeddings (LREs) and Linear Relational Concepts (LRCs) for LLMs in PyTorch☆11Aug 7, 2024Updated 2 years ago
- Academic Evaluation Modules☆11Updated this week
- [ACL 2024 Findings] Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruning☆13Sep 2, 2024Updated 2 years ago
- ☆12Feb 16, 2024Updated 2 years ago
- R code for cost-effectiveness model evaluating directly acting oral anticoagulants (DOACs) for prevention of stroke in atrial fibrillatio…☆12Jan 14, 2020Updated 6 years ago
- mcp wrapper for openai built-in tools☆12Mar 13, 2025Updated last year
- SeeGULL is a broad-coverage stereotype dataset in English containing stereotypes about identity groups spanning 178 countries across 8 di…☆38Sep 25, 2023Updated 2 years ago
- A framework for few-shot evaluation of language models.☆13,920Sep 1, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Inspect: A framework for large language model evaluations☆2,718Updated this week
- ☆575May 21, 2026Updated 3 months ago
- auto ticket reservation program (python)☆14Jan 28, 2020Updated 6 years ago
- Data-Driven Evaluation for LLM-Powered Applications☆517Aug 10, 2026Updated 3 weeks ago
- 🌊 ADHD-friendly productivity app. One next step at a time, flow-based workflows.☆17Feb 9, 2026Updated 6 months ago
- A Hands on series on developing LLM applications☆70Sep 28, 2024Updated last year
- Overall Equipment Effectiveness: Performant and Scalable End-to-End Equipment Monitoring☆13Jun 15, 2023Updated 3 years ago