Use Muon optimizer instead of AdamW.
☆49Mar 2, 2026Updated 5 months ago
Alternatives and similar repositories for muon-optimizer-guide
Users that are interested in muon-optimizer-guide are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Minimal and highly hackable implementation of Looped Transformers with GPT☆25Mar 8, 2026Updated 5 months ago
- Conformal Prediction + Federated Learning☆16Mar 16, 2024Updated 2 years ago
- A set of markdown files to point Claude towards to get an amazing mandarin tutor☆16May 23, 2026Updated 2 months ago
- Notes for CS294/194-196: Large Language Model Agents (Fall 2024, UC Berkeley), summarizing 12 lectures on LLM fundamentals, reasoning, pl…☆18Jan 7, 2025Updated last year
- An efficient and scalable attention module designed to reduce memory usage and improve inference speed in large language models. Designe…☆26Jun 25, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"☆14Jun 21, 2026Updated last month
- ☆24Jan 30, 2025Updated last year
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆12Oct 31, 2024Updated last year
- Flax (JAX) implementation of Progressive Growing of GANs for Improved Quality, Stability, and Variation☆12May 24, 2021Updated 5 years ago
- optimize neuro-centric parameters instead of weights to solve RL tasks☆14Oct 2, 2023Updated 2 years ago
- Minimal open-source implementation of AlphaProof and HyperTree Proof Search.☆87May 13, 2026Updated 3 months ago
- ☆13May 10, 2024Updated 2 years ago
- ☆15Sep 29, 2022Updated 3 years ago
- Source code for our paper: "Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents".☆25Feb 20, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Benchmarking foundation Machine Learning Potentials with Lattice Thermal Conductivity from Anharmonic Phonons☆15Oct 30, 2024Updated last year
- A framework to meta-train transformers for causal ICL☆11Jul 15, 2026Updated last month
- 100M tokens. Infinite compute. Lowest val loss wins.☆523Jul 3, 2026Updated last month
- A Gentle Principled Introduction to Deep Reinforcement Learning☆19Apr 4, 2025Updated last year
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆28May 12, 2026Updated 3 months ago
- [Preprint] AdaVAE: Exploring Adaptive GPT-2s in VAEs for Language Modeling PyTorch Implementation☆38Oct 18, 2023Updated 2 years ago
- Mirror of the ASKCOSv2 Project with Chemhacktica-flavored changes☆25Jul 31, 2025Updated last year
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"☆15Sep 25, 2025Updated 10 months ago
- ☆15Apr 15, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Educational WIP☆73Feb 16, 2026Updated 5 months ago
- ☆28Jun 3, 2022Updated 4 years ago
- ☆22May 12, 2026Updated 3 months ago
- Utilities for PyTorch distributed☆26Feb 27, 2025Updated last year
- Train LLM from scratch for $5 USD - Research.☆210Feb 23, 2026Updated 5 months ago
- jQMC code implements two real-space ab initio quantum Monte Carlo (QMC) methods. Variatioinal Monte Carlo (VMC) and lattice regularized d…☆20Jul 7, 2026Updated last month
- Open-source Human Feedback Library☆11Oct 25, 2023Updated 2 years ago
- DLLM-Searcher has been accepted by SIGIR 2026! 🥳☆33Jan 23, 2026Updated 6 months ago
- Distilled version of LaGAT: a neuro-guided search for multi-agent pathfinding (MAPF)☆15Jul 18, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆19Mar 3, 2026Updated 5 months ago
- An automatic differentiation system for dense and sparse problems☆13Jan 16, 2025Updated last year
- CIFAR-10 speedrun: Trains to 94% accuracy in 1.98 seconds on a single NVIDIA A100 GPU.☆79Jul 30, 2026Updated 2 weeks ago
- Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding (ICLR-2026)☆16Mar 3, 2026Updated 5 months ago
- A fast, memory-efficient exact MaxSim kernel for late-interaction retrieval and reranking.☆26May 27, 2026Updated 2 months ago
- exBERT on Transformers🤗☆10Jun 14, 2021Updated 5 years ago
- A modified Ziggurat Algorithm for efficiently generating exponentially- and normally-distributed PseudoRandom Numbers (PRNs).☆13May 21, 2025Updated last year