Official implementation of 'A Large-Scale Exploration of mu-Transfer' (CoRR 2024)
☆31Jun 5, 2025Updated last year
Alternatives and similar repositories for mu_transformer
Users that are interested in mu_transformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Maximal Update Parametrization (μP) with Flax & Optax.☆16Dec 27, 2023Updated 2 years ago
- Minimal but scalable implementation of large language models in JAX☆34Nov 28, 2025Updated 8 months ago
- ☆19Jul 8, 2026Updated last month
- Don't just regulate gradients like in Muon, regulate the weights too☆32Jul 30, 2025Updated last year
- (EasyDel Former) is a utility library designed to simplify and enhance the development in JAX☆33Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Morphological segmentation of Galaxies☆10Dec 17, 2021Updated 4 years ago
- ☆25Dec 16, 2025Updated 8 months ago
- ☆16Oct 30, 2024Updated last year
- Unofficial JAX implementation of the SOAP optimizer (https://arxiv.org/abs/2409.11321)☆29Jul 21, 2026Updated 3 weeks ago
- Two implementations of ZeRO-1 optimizer sharding in JAX☆14Jun 11, 2023Updated 3 years ago
- A from-scratch neural network and transformers library, with speeds rivaling PyTorch☆10Mar 16, 2025Updated last year
- nanoGPT using Equinox☆15Mar 3, 2023Updated 3 years ago
- A library for unit scaling in PyTorch☆136Jul 11, 2025Updated last year
- implementation of https://arxiv.org/pdf/2312.09299☆21Jul 3, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- JAX Scalify: end-to-end scaled arithmetics☆18Oct 30, 2024Updated last year
- Schedule free optimiser implemented in JAX using Optimistix☆15May 29, 2024Updated 2 years ago
- some common Huggingface transformers in maximal update parametrization (µP)☆88Mar 14, 2022Updated 4 years ago
- Cookbook for Crafting Good Code☆56Mar 19, 2024Updated 2 years ago
- 4-bit Shampoo for Memory-Efficient Network Training (NeurIPS 2024)☆13Feb 13, 2025Updated last year
- Supercharge huggingface transformers with model parallelism.☆77Jul 23, 2025Updated last year
- Code for the paper "Function-Space Learning Rates"☆23Jun 3, 2025Updated last year
- Implemented stackless KDTree on GPU to accelerate ray tracing rendering algorithm. Hardware level optimizations for register spills local…☆21Oct 6, 2016Updated 9 years ago
- A collection of optimizers, some arcane others well known, for Flax.☆29Aug 6, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation of PSGD optimizer in JAX☆36Dec 31, 2024Updated last year
- A repo based on XiLin Li's PSGD repo that extends some of the experiments.☆14Oct 7, 2024Updated last year
- ☆13Apr 25, 2024Updated 2 years ago
- Minimal (400 LOC) implementation Maximum (multi-node, FSDP) GPT training☆132Apr 17, 2024Updated 2 years ago
- A fully trainable state space model (SSM)☆16Mar 18, 2025Updated last year
- Rust implementation of Surya☆66Mar 1, 2025Updated last year
- Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.☆19Jul 24, 2025Updated last year
- Bullseye Polytope Clean-Label Poisoning Attack☆19Nov 5, 2020Updated 5 years ago
- ☆18Aug 24, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A github action to setup a small SLURM cluster for testing purposes.☆14Jul 20, 2025Updated last year
- AI for a cure, a combination of Latent-GAN and VAE-JTNN to create 100% valid drug like molecules☆10Mar 16, 2020Updated 6 years ago
- Train to 94% on CIFAR-10 in 4.4 seconds on a single A100☆12Dec 30, 2023Updated 2 years ago
- A PyTorch implementation of perceptual loss using ConvNeXt feature extractors.☆30Nov 14, 2024Updated last year
- seqax = sequence modeling + JAX☆198Jul 23, 2025Updated last year
- Minimal Implimentation of VCRec (2024) for collapse provention.☆18Jan 28, 2025Updated last year
- Lightweight tools for quick and easy LLM demo's☆28Sep 22, 2024Updated last year