A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
β24,805Aug 15, 2024Updated 2 years ago
Alternatives and similar repositories for minGPT
Users that are interested in minGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The simplest, fastest repository for training/finetuning medium-sized GPTs.β62,126Nov 12, 2025Updated 9 months ago
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β164,109Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β42,939Updated this week
- Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.β31,286Updated this week
- Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and moreβ36,161Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Inference Llama 2 in one file of pure Cβ19,983Aug 6, 2024Updated 2 years ago
- Facebook AI Research Sequence-to-Sequence Toolkit written in Python.β32,240Sep 30, 2025Updated 10 months ago
- A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like APIβ17,137Aug 3, 2026Updated last week
- Code for the paper "Language Models are Unsupervised Multitask Learners"β25,024Aug 14, 2024Updated 2 years ago
- Google Researchβ38,540Updated this week
- Making large AI models cheaper, faster and more accessibleβ41,437Updated this week
- LLM training in simple, raw C/CUDAβ30,809Jun 26, 2025Updated last year
- Fast and memory-efficient exact attentionβ24,716Updated this week
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,189Jan 23, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,244Jul 17, 2024Updated 2 years ago
- Inference code for Llama modelsβ59,557Jan 26, 2025Updated last year
- Train transformer language models with reinforcement learning.β19,076Updated this week
- You like pytorch? You like micrograd? You love tinygrad! β€οΈβ33,452Updated this week
- A library for efficient similarity search and clustering of dense vectors.β40,744Updated this week
- Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)β9,573Jul 5, 2026Updated last month
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,513May 1, 2026Updated 3 months ago
- Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.β10,679Jul 1, 2024Updated 2 years ago
- π€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.β34,319Updated this week
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Neural Networks: Zero to Heroβ23,983Aug 18, 2024Updated last year
- The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights --β¦β37,066Updated this week
- CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an imageβ34,171Mar 25, 2026Updated 4 months ago
- Ongoing research training transformer models at scaleβ17,439Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMsβ89,112Updated this week
- Development repository for the Triton language and compilerβ19,950Updated this week
- Tensors and Dynamic neural networks in Python with strong GPU accelerationβ102,388Updated this week
- OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamicalβ¦β37,406Aug 17, 2024Updated last year
- The agent engineering platform.β144,261Updated this week
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- LLM inference in C/C++β123,999Updated this week
- Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.β43,520Updated this week
- LlamaIndex is the leading document agent and OCR platformβ51,658Updated this week
- Build and share delightful machine learning apps, all in Python. π Star to support our work!β43,365Updated this week
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,552Updated this week
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,819Updated this week
- RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable)β¦β14,663Updated this week