CIFAR-10 speedruns: 94% in 2.6 seconds and 96% in 27 seconds
☆386Nov 15, 2025Updated 9 months ago
Alternatives and similar repositories for cifar10-airbench
Users that are interested in cifar10-airbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CIFAR-10 speedrun: Trains to 94% accuracy in 1.98 seconds on a single NVIDIA A100 GPU.☆79Jul 30, 2026Updated 2 weeks ago
- NanoGPT (124M) in 90 seconds☆5,681Aug 9, 2026Updated last week
- Muon is an optimizer for hidden layers in neural networks☆2,787May 24, 2026Updated 2 months ago
- Implementation of PSGD optimizer in JAX☆36Dec 31, 2024Updated last year
- 🧱 Modula software package☆337Aug 18, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Efficient optimizers☆339Jul 25, 2026Updated 3 weeks ago
- Code for the paper "Function-Space Learning Rates"☆23Jun 3, 2025Updated last year
- Pytorch implementation of preconditioned stochastic gradient descent (Kron and affine preconditioner, low-rank approximation precondition…☆199May 30, 2026Updated 2 months ago
- supporting pytorch FSDP for optimizers☆84Dec 8, 2024Updated last year
- Train to 94% on CIFAR-10 in <6.3 seconds on a single A100. Or ~95.79% in ~110 seconds (or less!)☆1,310Dec 18, 2024Updated last year
- Schedule-Free Optimization in PyTorch☆2,322Jul 28, 2026Updated 3 weeks ago
- Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.☆19Jul 24, 2025Updated last year
- ☆70Apr 8, 2026Updated 4 months ago
- Custom triton kernels for training Karpathy's nanoGPT.☆19Oct 21, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Minimal (truly) muP implementation, consistent with TP4 and TP5 papers notation☆14Jan 2, 2026Updated 7 months ago
- Dion optimizer algorithm☆535Updated this week
- 100M tokens. Infinite compute. Lowest val loss wins.☆526Jul 3, 2026Updated last month
- ☆278Dec 2, 2024Updated last year
- ☆27May 3, 2024Updated 2 years ago
- A fully trainable state space model (SSM)☆16Mar 18, 2025Updated last year
- Flash-Muon: An Efficient Implementation of Muon Optimizer☆260Jun 15, 2025Updated last year
- ☆34Jan 25, 2024Updated 2 years ago
- An implementation of PSGD Kron second-order optimizer for PyTorch☆102Jul 24, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Minimalistic, extremely fast, and hackable researcher's toolbench for GPT models in 307 lines of code. Reaches <3.8 validation loss on wi…☆359Jul 29, 2024Updated 2 years ago
- Code for AAAI 2024 paper: CR-SAM: Curvature Regularized Sharpness-Aware Minimization☆12Nov 29, 2024Updated last year
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆201Jan 19, 2026Updated 6 months ago
- Data for "Datamodels: Predicting Predictions with Training Data"☆97May 25, 2023Updated 3 years ago
- Minimal implementation of scalable rectified flow transformers, based on SD3's approach☆643Jul 1, 2024Updated 2 years ago
- Open-source framework for the research and development of foundation models.☆1,267Updated this week
- Python library for argument and configuration management☆57Feb 7, 2023Updated 3 years ago
- ☆23Jan 23, 2024Updated 2 years ago
- ☆24Jun 18, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A comprehensive JAX/NNX library for diffusion and flow matching generative algorithms, featuring DiT (Diffusion Transformer) and its vari…☆154Oct 16, 2025Updated 10 months ago
- Don't just regulate gradients like in Muon, regulate the weights too☆32Jul 30, 2025Updated last year
- PyTorch linear operators for curvature matrices (Hessian, Fisher/GGN, KFAC, ...)☆71Jul 17, 2026Updated last month
- Cuda implemenation of flash-kmeans, 2x faster☆23May 8, 2026Updated 3 months ago
- Focused on fast experimentation and simplicity☆77Dec 24, 2024Updated last year
- Grams: Gradient Descent with Adaptive Momentum Scaling (ICLR 2025 Workshop)☆17Mar 6, 2025Updated last year
- WIP☆96Aug 13, 2024Updated 2 years ago