π§± Modula software package
β337Aug 18, 2025Updated 11 months ago
Alternatives and similar repositories for modula
Users that are interested in modula are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Don't just regulate gradients like in Muon, regulate the weights tooβ32Jul 30, 2025Updated 11 months ago
- A single-line modification to any (dualizer-based) optimizer that allows the optimizer to adapt to the scale of the gradients as they chaβ¦β19Jan 11, 2025Updated last year
- Code for the paper "Function-Space Learning Rates"β23Jun 3, 2025Updated last year
- Pytorch implementation of preconditioned stochastic gradient descent (Kron and affine preconditioner, low-rank approximation preconditionβ¦β198May 30, 2026Updated last month
- [Poster; ICLR 2026] [Oral; Neurips OPT2024] ΞΌLO: Compute-Efficient Meta-Generalization of Learned Optimizersβ16Apr 15, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Supporting code for the blog post on modular manifolds.β125Sep 26, 2025Updated 9 months ago
- β34Oct 4, 2024Updated last year
- Experiment of using Tangent to autodiff tritonβ82Jan 22, 2024Updated 2 years ago
- Efficient optimizersβ334Jul 11, 2026Updated last week
- A library for unit scaling in PyTorchβ134Jul 11, 2025Updated last year
- supporting pytorch FSDP for optimizersβ84Dec 8, 2024Updated last year
- An implementation of PSGD Kron second-order optimizer for PyTorchβ102Jul 24, 2025Updated 11 months ago
- WIPβ96Aug 13, 2024Updated last year
- β70Apr 8, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Maximal Update Parametrization (ΞΌP) with Flax & Optax.β16Dec 27, 2023Updated 2 years ago
- Simple implementation of muP, based on Spectral Condition for Feature Learning. The implementation is SGD only, dont use it for Adamβ88Jul 28, 2024Updated last year
- β63Oct 3, 2024Updated last year
- The simplest, fastest repository for training/finetuning medium-sized GPTs.β199Jan 19, 2026Updated 6 months ago
- Dion optimizer algorithmβ494Jul 12, 2026Updated last week
- Understand and test language model architectures on synthetic tasks.β277Mar 22, 2026Updated 3 months ago
- Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.β19Jul 24, 2025Updated 11 months ago
- β273Dec 2, 2024Updated last year
- Minimal (truly) muP implementation, consistent with TP4 and TP5 papers notationβ14Jan 2, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Compositional Linear Algebraβ517Aug 1, 2025Updated 11 months ago
- β13Mar 10, 2026Updated 4 months ago
- Code for implementing central flowsβ48Sep 5, 2025Updated 10 months ago
- CIFAR-10 speedruns: 94% in 2.6 seconds and 96% in 27 secondsβ383Nov 15, 2025Updated 8 months ago
- Tile primitives for speedy kernelsβ3,550Jul 13, 2026Updated last week
- β45Nov 1, 2025Updated 8 months ago
- Combining SOAP and MUONβ22Feb 11, 2025Updated last year
- Fine-Tuning Pre-trained Transformers into Decaying Fast Weightsβ20Oct 9, 2022Updated 3 years ago
- 4-bit Shampoo for Memory-Efficient Network Training (NeurIPS 2024)β13Feb 13, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Flash-Muon: An Efficient Implementation of Muon Optimizerβ257Jun 15, 2025Updated last year
- Schedule-Free Optimization in PyTorchβ2,314Jun 18, 2026Updated last month
- Official implementation of 'A Large-Scale Exploration of mu-Transfer' (CoRR 2024)β31Jun 5, 2025Updated last year
- β304Jul 15, 2024Updated 2 years ago
- Minimal but scalable implementation of large language models in JAXβ34Nov 28, 2025Updated 7 months ago
- Uncertainty quantification with PyTorchβ384Apr 1, 2026Updated 3 months ago
- LoRA-Ensemble: Efficient Uncertainty Modelling for Self-attention Networksβ55Mar 7, 2026Updated 4 months ago