Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.
☆19Jul 24, 2025Updated last year
Alternatives and similar repositories for llm-jax
Users that are interested in llm-jax are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Maximal Update Parametrization (μP) with Flax & Optax.☆16Dec 27, 2023Updated 2 years ago
- Automatically take good care of your preemptible TPUs☆37May 15, 2023Updated 3 years ago
- Minimal but scalable implementation of large language models in JAX☆34Nov 28, 2025Updated 9 months ago
- ☆19Updated this week
- ☆24Jun 18, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- a Jax/Flax inference code of StarCoder☆12Jun 12, 2023Updated 3 years ago
- ☆32Nov 4, 2024Updated last year
- Tiny AutoEncoder for Stable Diffusion Videos☆37Oct 5, 2024Updated last year
- ☆22Dec 15, 2023Updated 2 years ago
- ☆23Jan 5, 2025Updated last year
- supporting pytorch FSDP for optimizers☆84Dec 8, 2024Updated last year
- ☆21Nov 18, 2024Updated last year
- Jax/Flax rewrite of Karpathy's nanoGPT☆66Feb 15, 2023Updated 3 years ago
- A flexible and efficient implementation of Flash Attention 2.0 for JAX, supporting multiple backends (GPU/TPU/CPU) and platforms (Triton/…☆34Mar 4, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- implementation of https://arxiv.org/pdf/2312.09299☆21Jul 3, 2024Updated 2 years ago
- (EasyDel Former) is a utility library designed to simplify and enhance the development in JAX☆33Aug 15, 2026Updated 3 weeks ago
- Pytorch implementation of preconditioned stochastic gradient descent (Kron and affine preconditioner, low-rank approximation precondition…☆205May 30, 2026Updated 3 months ago
- 4-bit Shampoo for Memory-Efficient Network Training (NeurIPS 2024)☆13Feb 13, 2025Updated last year
- Efficient optimizers☆340Aug 29, 2026Updated last week
- ☆19Dec 4, 2025Updated 9 months ago
- A repo based on XiLin Li's PSGD repo that extends some of the experiments.☆14Oct 7, 2024Updated last year
- ☆13Apr 25, 2024Updated 2 years ago
- A set of Python scripts that makes your experience on TPU better☆56Sep 18, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- CIFAR10 ResNets implemented in JAX+Flax☆12Apr 6, 2022Updated 4 years ago
- Tool for generating pictures using mathematical formulas.☆44Nov 8, 2021Updated 4 years ago
- ☆23Jan 23, 2024Updated 2 years ago
- KANs and MLPs☆12Jun 7, 2024Updated 2 years ago
- A fully trainable state space model (SSM)☆16Mar 18, 2025Updated last year
- Train to 94% on CIFAR-10 in 4.4 seconds on a single A100☆12Dec 30, 2023Updated 2 years ago
- A simple library for scaling up JAX programs☆149Nov 4, 2025Updated 10 months ago
- LLM shell and document interogator☆14Jul 24, 2023Updated 3 years ago
- some common Huggingface transformers in maximal update parametrization (µP)☆88Mar 14, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An implementation of DecorrelatedBN by tensorflow☆13Jun 30, 2022Updated 4 years ago
- An implementation of PSGD Kron second-order optimizer for PyTorch☆117Jul 24, 2025Updated last year
- LLM training in simple, raw C/CUDA☆15Dec 5, 2024Updated last year
- ☆10Feb 12, 2024Updated 2 years ago
- ☆18Aug 24, 2024Updated 2 years ago
- Repository for Skill Set Optimization☆14Jul 26, 2024Updated 2 years ago
- ☆21Sep 6, 2021Updated 5 years ago