Benchmarking Optimizers for LLM Pretraining
☆60May 3, 2026Updated 2 months ago
Alternatives and similar repositories for llm-optimizer-benchmark
Users that are interested in llm-optimizer-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ScalingOpt - Optimization Community☆104Jun 1, 2026Updated last month
- ☆70Apr 8, 2026Updated 3 months ago
- [ICML 2026] Memory-Efficient LLM Pretraining via Minimalist Optimizer Design☆21May 26, 2026Updated last month
- Spectral Sphere Optimizer☆130Mar 23, 2026Updated 3 months ago
- ☆14Mar 2, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- nanoGPT-like codebase for LLM training☆118Nov 7, 2025Updated 8 months ago
- Combining SOAP and MUON☆22Feb 11, 2025Updated last year
- The Newton-Muon optimizer☆30Jun 5, 2026Updated last month
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated last month
- Code for "What really matters in matrix-whitening optimizers?"☆25Oct 31, 2025Updated 8 months ago
- ☆273Dec 2, 2024Updated last year
- Aurora optimizer release☆150Updated this week
- ☆22Jul 13, 2026Updated last week
- ☆13May 4, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Grams: Gradient Descent with Adaptive Momentum Scaling (ICLR 2025 Workshop)☆17Mar 6, 2025Updated last year
- Fast Polar Decomposition for Muon☆166Jul 2, 2026Updated 2 weeks ago
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆199Jan 19, 2026Updated 6 months ago
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"☆15Sep 25, 2025Updated 9 months ago
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆12Oct 31, 2024Updated last year
- Subset-Norm and Subset-Momentum. This repo is built on top of https://github.com/jiaweizzhao/GaLore.☆19Jul 9, 2025Updated last year
- Official repository of PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective☆26Jun 13, 2026Updated last month
- ☆14Dec 13, 2024Updated last year
- ☆75Jun 23, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Computing various measures and generalization bounds on convolutional and fully connected networks☆35Dec 13, 2018Updated 7 years ago
- ☆18Aug 19, 2024Updated last year
- nanoGPT using Equinox☆15Mar 3, 2023Updated 3 years ago
- Pytorch code for experiments on Linear Transformers☆24Jan 12, 2024Updated 2 years ago
- Implementation for POET and POET-X for LLM pretraining☆38Jun 9, 2026Updated last month
- torch implementation of diloco☆24Updated this week
- The AdEMAMix Optimizer: Better, Faster, Older.☆188Sep 12, 2024Updated last year
- Code for reproducing the results from "CrAM: A Compression-Aware Minimizer" accepted at ICLR 2023☆10Mar 1, 2023Updated 3 years ago
- Accelerated Bregman Proximal Gradient Methods☆29Jun 12, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Implementation of Gradient Information Optimization (GIO) for effective and scalable training data selection☆14Jun 22, 2023Updated 3 years ago
- ☆63Oct 3, 2024Updated last year
- [EMNLP 2023, Main Conference] Sparse Low-rank Adaptation of Pre-trained Language Models☆87Mar 5, 2024Updated 2 years ago
- ☆26Feb 20, 2026Updated 5 months ago
- Pytorch routines for (Ker)nel (Mac)hines☆12Oct 10, 2025Updated 9 months ago
- Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations☆31Jul 2, 2026Updated 2 weeks ago
- [ICLR 2026] When it comes to optimizers, it's always better to be safe than sorry☆417Sep 26, 2025Updated 9 months ago