This repository contains code for the MicroAdam paper.
β21Dec 14, 2024Updated last year
Alternatives and similar repositories for MicroAdam
Users that are interested in MicroAdam are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β14May 4, 2026Updated 3 months ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. π The official implementation of https://arxβ¦β31Feb 17, 2025Updated last year
- Resources regarding evML (edge verified machine learning)β24Jan 4, 2025Updated last year
- SGLang Kernel Wheel Indexβ24Updated this week
- β27Aug 25, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code for "RSQ: Learning from Important Tokens Leads to Better Quantized LLMs"β23Mar 25, 2026Updated 4 months ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspaceβ19Oct 21, 2024Updated last year
- [ICML2024 Spotlight] Fine-Tuning Pre-trained Large Language Models Sparselyβ24Jun 26, 2024Updated 2 years ago
- Repository for Sparse Finetuning of LLMs via modified version of the MosaicML llmfoundryβ43Jan 15, 2024Updated 2 years ago
- β14Nov 3, 2025Updated 9 months ago
- Official Code of The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural Networks[ICML2022]β16Sep 20, 2022Updated 3 years ago
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learningβ10Apr 28, 2023Updated 3 years ago
- Physics Master is a model fine-tuned from llama3-8B-Instruct. It can answer your physics question!β16Aug 24, 2024Updated last year
- [ICML 2024] Official Implementation of SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocksβ43Feb 4, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation of "Can Test-Time Scaling Improve World Foundation Model?"β15Jul 12, 2025Updated last year
- Pytorch distributed backend extension with compression supportβ17Mar 24, 2025Updated last year
- Training with Block Minifloat number representationβ18May 2, 2021Updated 5 years ago
- β33Nov 11, 2024Updated last year
- CLI tool designed to manage MCI (Model Context Interface) schemas and dynamically run MCP servers using defined MCI toolsetsβ16Nov 12, 2025Updated 8 months ago
- Using LogMel Spectrograms to detect cough in audio samples using a CNNβ16Jun 18, 2020Updated 6 years ago
- Boosting 4-bit inference kernels with 2:4 Sparsityβ100Sep 4, 2024Updated last year
- A PyTorch Implementation of Neural Turing Machineβ14Jul 24, 2020Updated 6 years ago
- β61Jun 10, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Artifact evaluation for HPCA'24 paper Lightening-Transformer: A Dynamically-operated Optically-interconnected Photonic Transformer Acceleβ¦β11Mar 3, 2024Updated 2 years ago
- β17Apr 7, 2025Updated last year
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"β93Oct 30, 2024Updated last year
- QLoRA: Efficient Finetuning of Quantized LLMsβ11Jul 22, 2023Updated 3 years ago
- Inducing Point Operator Transformer: A Flexible and Scalable Architecture for Solving PDEs (AAAI 2024)β14Jul 30, 2024Updated 2 years ago
- Code for ICML 2022 paper "SPDY: Accurate Pruning with Speedup Guarantees"β20May 3, 2023Updated 3 years ago
- Implementation of RankE: End-to-End Discrete Text-to-Image Post-Training via Rank-Consistent Alignmentβ22May 27, 2026Updated 2 months ago
- easy exllama interface w/ automation & evalsβ17Aug 2, 2026Updated last week
- β10Jun 30, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- new optimizerβ20Aug 4, 2024Updated 2 years ago
- [ICLR'25] R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inferenceβ21Apr 28, 2025Updated last year
- Testing KAN-based text generation GPT modelsβ19May 6, 2024Updated 2 years ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"β17Jun 30, 2025Updated last year
- β11Mar 23, 2022Updated 4 years ago
- β15Sep 24, 2023Updated 2 years ago