☆281Dec 2, 2024Updated last year
Alternatives and similar repositories for SOAP
Users that are interested in SOAP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Efficient optimizers☆340Sep 9, 2026Updated 2 weeks ago
- ☆70Aug 28, 2026Updated 3 weeks ago
- Combining SOAP and MUON☆25Feb 11, 2025Updated last year
- ☆72Sep 11, 2026Updated last week
- For optimization algorithm research and development.☆580Sep 14, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for "What really matters in matrix-whitening optimizers?"☆25Oct 31, 2025Updated 10 months ago
- 🧱 Modula software package☆342Aug 18, 2025Updated last year
- Muon is an optimizer for hidden layers in neural networks☆2,853May 24, 2026Updated 3 months ago
- ☆36Aug 14, 2026Updated last month
- Grams: Gradient Descent with Adaptive Momentum Scaling (ICLR 2025 Workshop)☆17Mar 6, 2025Updated last year
- ☆14Mar 2, 2025Updated last year
- Schedule-Free Optimization in PyTorch☆2,323Jul 28, 2026Updated last month
- Pytorch implementation of preconditioned stochastic gradient descent (Kron and affine preconditioner, low-rank approximation precondition…☆206May 30, 2026Updated 3 months ago
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"☆94Oct 30, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Dion optimizer algorithm☆547Updated this week
- Unofficial JAX implementation of the SOAP optimizer (https://arxiv.org/abs/2409.11321)☆29Jul 21, 2026Updated 2 months ago
- Spectral Sphere Optimizer☆133Mar 23, 2026Updated 6 months ago
- Code for Adam-mini: Use Fewer Learning Rates To Gain More https://arxiv.org/abs/2406.16793☆459May 13, 2025Updated last year
- Code for the paper: Why Transformers Need Adam: A Hessian Perspective☆65Mar 11, 2025Updated last year
- ☆34Mar 14, 2025Updated last year
- An implementation of PSGD Kron second-order optimizer for PyTorch☆119Jul 24, 2025Updated last year
- Benchmarking Optimizers for LLM Pretraining☆61May 3, 2026Updated 4 months ago
- Experiments on the impact of depth in transformers and SSMs.☆47Oct 23, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- WIP☆96Aug 13, 2024Updated 2 years ago
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆203Jan 19, 2026Updated 8 months ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆69Jul 30, 2026Updated last month
- Official Implementation of "ADOPT: Modified Adam Can Converge with Any β2 with the Optimal Rate"☆438Dec 12, 2024Updated last year
- ☆14May 4, 2026Updated 4 months ago
- Focused on fast experimentation and simplicity☆77Dec 24, 2024Updated last year
- ☆10Jun 27, 2024Updated 2 years ago
- ☆19Dec 4, 2025Updated 9 months ago
- ☆54May 20, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for "Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining"☆30Oct 14, 2025Updated 11 months ago
- [ICLR 2026] When it comes to optimizers, it's always better to be safe than sorry☆421Sep 26, 2025Updated 11 months ago
- Triton kernels for dynamic causal short convolutions.☆30Jun 4, 2026Updated 3 months ago
- An efficient implementation of the NSA (Native Sparse Attention) kernel☆135Jun 24, 2025Updated last year
- ☆20Feb 2, 2026Updated 7 months ago
- Supporting code for the blog post on modular manifolds.☆130Sep 26, 2025Updated 11 months ago
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated 3 months ago