A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
β24,732Aug 15, 2024Updated last year
Alternatives and similar repositories for minGPT
Users that are interested in minGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The simplest, fastest repository for training/finetuning medium-sized GPTs.β61,512Nov 12, 2025Updated 8 months ago
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β162,967Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β42,802Updated this week
- Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.β31,250Updated this week
- Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and moreβ36,050Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Inference Llama 2 in one file of pure Cβ19,772Aug 6, 2024Updated last year
- Facebook AI Research Sequence-to-Sequence Toolkit written in Python.β32,253Sep 30, 2025Updated 9 months ago
- A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like APIβ16,845Aug 8, 2024Updated last year
- Code for the paper "Language Models are Unsupervised Multitask Learners"β25,022Aug 14, 2024Updated last year
- Google Researchβ38,430Updated this week
- Making large AI models cheaper, faster and more accessibleβ41,427Jul 13, 2026Updated last week
- LLM training in simple, raw C/CUDAβ30,640Jun 26, 2025Updated last year
- Fast and memory-efficient exact attentionβ24,531Updated this week
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,171Jan 23, 2026Updated 6 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,252Jul 17, 2024Updated 2 years ago
- Inference code for Llama modelsβ59,530Jan 26, 2025Updated last year
- Train transformer language models with reinforcement learning.β18,927Updated this week
- You like pytorch? You like micrograd? You love tinygrad! β€οΈβ33,349Updated this week
- A library for efficient similarity search and clustering of dense vectors.β40,583Updated this week
- Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)β9,557Jul 5, 2026Updated 3 weeks ago
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,503May 1, 2026Updated 2 months ago
- Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.β10,640Jul 1, 2024Updated 2 years ago
- π€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.β34,151Updated this week
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Neural Networks: Zero to Heroβ23,665Aug 18, 2024Updated last year
- The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights --β¦β37,013Jul 16, 2026Updated last week
- CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an imageβ34,069Mar 25, 2026Updated 4 months ago
- Ongoing research training transformer models at scaleβ17,212Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMsβ87,138Updated this week
- Development repository for the Triton language and compilerβ19,782Updated this week
- Tensors and Dynamic neural networks in Python with strong GPU accelerationβ101,947Updated this week
- OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamicalβ¦β37,386Aug 17, 2024Updated last year
- The agent engineering platform.β142,576Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LLM inference in C/C++β121,564Updated this week
- Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.β43,350Updated this week
- LlamaIndex is the leading document agent and OCR platformβ51,086Updated this week
- Build and share delightful machine learning apps, all in Python. π Star to support our work!β43,206Updated this week
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,450Updated this week
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,794Updated this week
- RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable)β¦β14,639Updated this week