Curated list of datasets and tools for post-training.
☆4,762Apr 29, 2026Updated 4 months ago
Alternatives and similar repositories for llm-datasets
Users that are interested in llm-datasets are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verifi…☆3,385Updated this week
- A quick guide (especially) for trending instruction finetuning datasets☆3,412Nov 28, 2023Updated 2 years ago
- A framework for few-shot evaluation of language models.☆13,879Updated this week
- Tools for merging pretrained large language models.☆7,337Jun 17, 2026Updated 2 months ago
- Go ahead and axolotl questions☆12,436Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.☆82,246Feb 5, 2026Updated 6 months ago
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends☆2,534Aug 11, 2026Updated 3 weeks ago
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.☆75,555Updated this week
- Robust recipes to align language models with human and AI preferences☆5,670May 26, 2026Updated 3 months ago
- Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard a…☆2,143Dec 3, 2025Updated 9 months ago
- Minimalistic large language model 3D-parallelism training☆2,807May 26, 2026Updated 3 months ago
- A course on aligning smol models.☆6,734Aug 20, 2026Updated 2 weeks ago
- 20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.☆13,648Updated this week
- Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)☆74,553Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Train transformer language models with reinforcement learning.☆19,211Updated this week
- Automatically evaluate your LLMs in Google Colab☆697May 7, 2024Updated 2 years ago
- Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We als…☆18,558May 19, 2026Updated 3 months ago
- AllenAI's post-training codebase☆3,857Updated this week
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.☆3,317Aug 13, 2026Updated 3 weeks ago
- DSPy: The framework for programming—not prompting—language models☆37,750Updated this week
- The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices☆5,311Apr 22, 2026Updated 4 months ago
- Fast Multimodal Semantic Deduplication & Filtering☆963May 24, 2026Updated 3 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆90,889Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The LLM Evaluation Framework☆18,081Updated this week
- A reading list on LLM based Synthetic Data Generation 🔥☆1,552Jun 5, 2025Updated last year
- PyTorch native post-training library☆5,802Updated this week
- Structured Outputs☆15,743Updated this week
- Fully open reproduction of DeepSeek-R1☆26,448Apr 2, 2026Updated 5 months ago
- Implement a ChatGPT-like LLM in PyTorch from scratch, step by step☆104,274Updated this week
- This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed not…☆29,352Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,258Updated this week
- awesome synthetic (text) datasets☆339Aug 2, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Summarize existing representative LLMs text datasets.☆1,480Mar 11, 2026Updated 5 months ago
- SGLang is a high-performance serving framework for large language models and multimodal models.☆33,856Updated this week
- Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets☆5,095Updated this week
- llama3 implementation one matrix multiplication at a time☆15,220May 23, 2024Updated 2 years ago
- Optimizing inference proxy for LLMs☆4,261Jul 18, 2026Updated last month
- Democratizing Reinforcement Learning for LLMs☆5,814Aug 24, 2026Updated last week
- ☆3,094Jun 16, 2026Updated 2 months ago