A simple reference implementation of the single-worker MuLoCo optimizer in Jax & PyTorch. MuLoCo-1 has been shown to outperfrom Muon and have larger critical batch sizes.
☆32Feb 27, 2026Updated 5 months ago
Alternatives and similar repositories for muloco-1
Users that are interested in muloco-1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An efficient implementation of learned optimizers in PyTorch☆61Jun 23, 2026Updated last month
- Codebase for " Reducing Representation Drift in Online Continual Learning"☆14Jun 8, 2021Updated 5 years ago
- CCLoco: Scaling Up Top-K Error Feedback with Local Optimizers☆27Aug 22, 2025Updated 11 months ago
- [Poster; ICLR 2026] [Oral; Neurips OPT2024] μLO: Compute-Efficient Meta-Generalization of Learned Optimizers☆16Apr 15, 2026Updated 3 months ago
- ACCO: An optimization algorithm for sharded distributed LLM training.☆13May 22, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [Oral; CVPR'22] Parametric Scattering Networks☆29Nov 10, 2025Updated 8 months ago
- ISMIR 2021: Curriculum Learning for Imbalanced Classification in Large Vocabulary Automatic Chord Recognition☆10Nov 8, 2021Updated 4 years ago
- A framework for evaluating LLMs in Atari games☆15Apr 21, 2025Updated last year
- ☆20Oct 21, 2022Updated 3 years ago
- [WACV'24] Object Re-Identification from Point Clouds☆20Jan 16, 2026Updated 6 months ago
- [ICML 2023] Decentralized SGD and Average-direction SAM are Asymptotically Equivalent☆20Dec 4, 2023Updated 2 years ago
- A comprehensive toolkit for streamlining data editing, search, and inspection for large-scale language model training and interpretabilit…☆21Oct 30, 2025Updated 9 months ago
- ScatNetLight for fast classifications of signals via Scattering Networks☆15Apr 20, 2017Updated 9 years ago
- Code repository for the paper "Geometric Scattering for Graph Data Analysis"☆14Aug 26, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Training hybrid models for dummies.☆31Nov 1, 2025Updated 9 months ago
- ☆15Sep 29, 2022Updated 3 years ago
- ☆80Oct 19, 2017Updated 8 years ago
- [TMLR 2024] Revisiting Random Weight Perturbation for Efficiently Improving Generalization☆12Oct 18, 2024Updated last year
- A port of muP to JAX/Haiku☆25Oct 23, 2022Updated 3 years ago
- Thinker project☆16Sep 4, 2024Updated last year
- An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.☆21Nov 28, 2022Updated 3 years ago
- Official repository for ICCV 2023: Get the Best of Both Worlds: Improving Accuracy and Transferability by Grassmann Class☆13Oct 16, 2023Updated 2 years ago
- ☆88Jun 16, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Apr 7, 2022Updated 4 years ago
- A data-driven, fast driving simulator for multi-agent coordination under partial observability.☆39Aug 29, 2024Updated last year
- Privacy & Security Principles, Documents and Testing☆11Jul 28, 2020Updated 6 years ago
- Learnable Graph Discovery☆10May 17, 2019Updated 7 years ago
- ☆20Apr 16, 2025Updated last year
- Project showing how to develop NKI kernels for Llama 3.2 1B inference☆21May 29, 2025Updated last year
- ☆13Feb 5, 2024Updated 2 years ago
- The official implementation of the "Hypernetwork approach to generating point clouds" paper☆27Dec 17, 2023Updated 2 years ago
- An implementation of the Equivariant Graph Neural Network (EGNN) layer type for DGL-PyTorch.☆15Dec 27, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A curated list of awesome tools and resources used at Qurit.☆11Dec 31, 2020Updated 5 years ago
- Courses, assignments & solutions for Stanford CS109☆15Feb 20, 2019Updated 7 years ago
- Binary Sentiment Analysis on Amazon Reviews by fine tuning pre trained XLNet☆14May 4, 2020Updated 6 years ago
- Memory Replay with Data Compression (ICLR 2022)☆16Sep 26, 2023Updated 2 years ago
- ☆19Apr 16, 2022Updated 4 years ago
- A simple but well-performing "single-hop" visual attention model for the GQA dataset☆20Aug 8, 2019Updated 7 years ago
- ☆17Dec 11, 2022Updated 3 years ago