Automatically evaluate your LLMs in Google Colab
☆697May 7, 2024Updated 2 years ago
Alternatives and similar repositories for llm-autoeval
Users that are interested in llm-autoeval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tools for merging pretrained large language models.☆7,288Jun 17, 2026Updated last month
- Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verifi…☆3,367Updated this week
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends☆2,517Updated this week
- Go ahead and axolotl questions☆12,349Updated this week
- Robust recipes to align language models with human and AI preferences☆5,659May 26, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This is our own implementation of 'Layer Selective Rank Reduction'☆240May 26, 2024Updated 2 years ago
- The official evaluation suite and dynamic data release for MixEval.☆254Nov 10, 2024Updated last year
- ☆126Dec 18, 2024Updated last year
- Curated list of datasets and tools for post-training.☆4,734Apr 29, 2026Updated 3 months ago
- A framework for few-shot evaluation of language models.☆13,616Updated this week
- ☆56Nov 6, 2024Updated last year
- Accelerate your Hugging Face Transformers 7.6-9x. Native to Hugging Face and PyTorch.☆684Aug 22, 2024Updated last year
- Fast & more realistic evaluation of chat language models. Includes leaderboard.☆188Dec 23, 2023Updated 2 years ago
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.☆3,262Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Just a bunch of benchmark logs for different LLMs☆130Jul 28, 2024Updated 2 years ago
- A bagel, with everything.☆326Apr 11, 2024Updated 2 years ago
- Evaluating LLMs with CommonGen-Lite☆95Mar 21, 2024Updated 2 years ago
- Curate better data for LLMs☆1,072Mar 19, 2024Updated 2 years ago
- Official repository for ORPO☆479May 31, 2024Updated 2 years ago
- Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs☆3,823May 28, 2026Updated 2 months ago
- Let's build better datasets, together!☆275Jun 9, 2026Updated 2 months ago
- 20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.☆13,613Updated this week
- Minimalistic large language model 3D-parallelism training