Minimalistic, extremely fast, and hackable researcher's toolbench for GPT models in 307 lines of code. Reaches <3.8 validation loss on wikitext-103 on a single A100 in <100 seconds. Scales to larger models with one parameter change (feature currently in alpha).
☆360Jul 29, 2024Updated 2 years ago
Alternatives and similar repositories for hlb-gpt
Users that are interested in hlb-gpt are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train to 94% on CIFAR-10 in <6.3 seconds on a single A100. Or ~95.79% in ~110 seconds (or less!)☆1,310Dec 18, 2024Updated last year
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆203Jan 19, 2026Updated 7 months ago
- ☆144Mar 31, 2023Updated 3 years ago
- Minimal (400 LOC) implementation Maximum (multi-node, FSDP) GPT training☆132Apr 17, 2024Updated 2 years ago
- It's a baby compiler. (Lean btw.)☆16May 19, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Collection of autoregressive model implementation☆85Aug 26, 2026Updated last week
- implementation of https://arxiv.org/pdf/2312.09299☆21Jul 3, 2024Updated 2 years ago
- Demonstration that finetuning RoPE model on larger sequences than the pre-trained model adapts the model context limit☆62Jun 21, 2023Updated 3 years ago
- Official implementation of 'A Large-Scale Exploration of mu-Transfer' (CoRR 2024)☆31Jun 5, 2025Updated last year
- Minimalistic, hackable PyTorch implementation of SimSiam in ~400 lines. Achieves good performance on ImageNet with ResNet50. Features dis…☆22Nov 25, 2024Updated last year
- Schedule-Free Optimization in PyTorch☆2,324Jul 28, 2026Updated last month
- Fast & Simple repository for pre-training and fine-tuning T5-style models☆1,021Aug 21, 2024Updated 2 years ago
- WIP☆96Aug 13, 2024Updated 2 years ago
- ☆125May 28, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Full finetuning of large language models without large memory requirements☆92Sep 22, 2025Updated 11 months ago
- CIFAR-10 speedruns: 94% in 2.6 seconds and 96% in 27 seconds☆389Nov 15, 2025Updated 9 months ago
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆714Jan 26, 2026Updated 7 months ago
- Low-Rank adapter extraction for fine-tuned transformers models☆181May 2, 2024Updated 2 years ago
- supporting pytorch FSDP for optimizers☆84Dec 8, 2024Updated last year
- an implementation of Self-Extend, to expand the context window via grouped attention☆118Jan 7, 2024Updated 2 years ago
- ☆306Jul 15, 2024Updated 2 years ago
- NanoGPT (124M) in 90 seconds☆5,735Aug 9, 2026Updated 3 weeks ago
- Entropy Based Sampling and Parallel CoT Decoding☆3,432Nov 13, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.☆604Aug 14, 2026Updated 2 weeks ago
- Cramming the training of a (BERT-type) language model into limited compute.☆1,366Jun 13, 2024Updated 2 years ago
- Code for the paper "Function-Space Learning Rates"☆23Jun 3, 2025Updated last year
- ☆62Mar 4, 2022Updated 4 years ago
- Efficient optimizers☆340Updated this week
- ☆314Jun 21, 2024Updated 2 years ago
- ☆54May 20, 2024Updated 2 years ago
- Just a bunch of benchmark logs for different LLMs☆130Jul 28, 2024Updated 2 years ago
- High-performance tokenized language data-loader for Python C++ extension☆15Jul 22, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for the paper "QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models".☆278Nov 3, 2023Updated 2 years ago
- 🚀 Efficiently (pre)training foundation models with native PyTorch features, including FSDP for training and SDPA implementation of Flash…☆288Nov 24, 2025Updated 9 months ago
- ☆14Oct 31, 2023Updated 2 years ago
- Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.☆19Jul 24, 2025Updated last year
- [WIP] Transformer to embed Danbooru labelsets☆13Mar 31, 2024Updated 2 years ago
- PyTorch interface for TrueGrad Optimizers☆43Aug 8, 2023Updated 3 years ago
- ☆13Jun 18, 2024Updated 2 years ago