Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
☆2,865Jul 1, 2026Updated 3 weeks ago
Alternatives and similar repositories for helm
Users that are interested in helm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A framework for few-shot evaluation of language models.☆13,452Jul 13, 2026Updated 2 weeks ago
- Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models☆3,250Jul 19, 2024Updated 2 years ago
- Benchmarking large language models' complex reasoning ability with chain-of-thought prompting☆2,776Aug 4, 2024Updated last year
- Measuring Massive Multitask Language Understanding | ICLR 2021☆1,605May 28, 2023Updated 3 years ago
- An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.☆2,008Aug 9, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.☆19,047Apr 14, 2026Updated 3 months ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)☆4,752Jan 8, 2024Updated 2 years ago
- Train transformer language models with reinforcement learning.☆18,953Updated this week
- 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.☆21,460Updated this week
- Toolkit for creating, sharing and using natural language prompts.☆3,027Oct 23, 2023Updated 2 years ago
- Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"☆1,853Jun 17, 2025Updated last year
- Aligning pretrained language models with instruction data generated by themselves.☆4,607Mar 27, 2023Updated 3 years ago
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends☆2,499Jun 29, 2026Updated last month
- OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, …☆7,246Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆4,586Apr 22, 2026Updated 3 months ago
- General technology for enabling AI capabilities w/ LLMs and MLLMs☆4,450Updated this week
- Robust recipes to align language models with human and AI preferences☆5,651May 26, 2026Updated 2 months ago
- ☆1,566Jul 2, 2026Updated 3 weeks ago
- The RedPajama-Data repository contains code for preparing large datasets for training large language models.☆4,975Jun 3, 2026Updated last month
- Ongoing research training transformer models at scale☆17,252Updated this week
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.☆39,508May 1, 2026Updated 2 months ago
- Code and documentation to train Stanford's Alpaca models, and generate the data.☆30,243Jul 17, 2024Updated 2 years ago
- Fast and memory-efficient exact attention☆24,568Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- AllenAI's post-training codebase☆3,813Updated this week
- Accessible large language models via k-bit quantization for PyTorch.☆8,369Updated this week
- This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on held-out tasks.☆552Mar 10, 2024Updated 2 years ago
- A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)☆3,614Feb 8, 2026Updated 5 months ago
- Tools for merging pretrained large language models.☆7,265Jun 17, 2026Updated last month
- ☆774Jun 13, 2024Updated 2 years ago
- The hub for EleutherAI's work on interpretability and learning dynamics☆2,867Nov 15, 2025Updated 8 months ago
- The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".☆1,610Apr 17, 2026Updated 3 months ago
- LLM training code for Databricks foundation models☆4,432Mar 25, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Minimalistic large language model 3D-parallelism training☆2,768May 26, 2026Updated 2 months ago
- 800,000 step-level correctness labels on LLM solutions to MATH problems☆2,152Jun 1, 2023Updated 3 years ago
- A modular RL library to fine-tune language models to human preferences☆2,393Mar 1, 2024Updated 2 years ago
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities☆22,173Jan 23, 2026Updated 6 months ago
- Large Language Model Text Generation Inference☆10,884Mar 21, 2026Updated 4 months ago
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.☆3,239Jul 22, 2026Updated last week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.☆42,830Updated this week