Code to generate figures of paper "When do spectral gradient updates help in deep learning?"
☆16Dec 3, 2025Updated 8 months ago
Alternatives and similar repositories for specgd
Users that are interested in specgd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34Aug 14, 2026Updated 2 weeks ago
- Spectral Sphere Optimizer☆133Mar 23, 2026Updated 5 months ago
- ☆70Updated this week
- A single-line modification to any (dualizer-based) optimizer that allows the optimizer to adapt to the scale of the gradients as they cha…☆19Jan 11, 2025Updated last year
- ☆16Dec 11, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Implementation of "RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm"☆18Apr 11, 2025Updated last year
- Github Repository for the HOI4 ULTRA Project.☆11Updated this week
- ☆23Jun 12, 2025Updated last year
- A toy eval suite for tracing generalization dynamics of LM pre-training☆21May 19, 2026Updated 3 months ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆31Aug 19, 2025Updated last year
- Code for ICLR 2023 Harnessing Out-Of-Distribution Examples via Augmenting Content and Style☆13Jul 3, 2023Updated 3 years ago
- Combining SOAP and MUON☆25Feb 11, 2025Updated last year
- [ICML 2025] Generative Modeling Reinvents Supervised Learning: Label Repurposing with Predictive Consistency Learning☆15Jul 14, 2025Updated last year
- ☆10Mar 23, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆26Feb 20, 2026Updated 6 months ago
- Code for paper Almost-Orthogonal Layers for Efficient General-Purpose Lipschitz Networks☆13Aug 9, 2022Updated 4 years ago
- Official github page for the paper "Evaluating Deep Unlearning in Large Language Model"☆14Apr 25, 2025Updated last year
- ☆52Mar 9, 2026Updated 5 months ago
- GeoZarr extension for OpenLayers☆12Jun 27, 2024Updated 2 years ago
- Don't just regulate gradients like in Muon, regulate the weights too☆32Jul 30, 2025Updated last year
- ☆26Jun 29, 2025Updated last year
- ☆31Dec 31, 2021Updated 4 years ago
- [NeurIPS 2025] Official implementation for our paper "Scaling Diffusion Transformers Efficiently via μP".☆100Nov 2, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Short RL☆19Apr 16, 2026Updated 4 months ago
- [ACL 2025] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models☆39Nov 4, 2025Updated 9 months ago
- ☆24Dec 11, 2024Updated last year
- Official Code for ACL 2023 Outstanding Paper: World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Languag…☆33Oct 20, 2023Updated 2 years ago
- My attempt to improve the speed of the newton schulz algorithm, starting from the dion implementation.☆42Apr 30, 2026Updated 4 months ago
- toy reproduction of Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts☆31Sep 1, 2024Updated last year
- Code for the paper "Function-Space Learning Rates"☆23Jun 3, 2025Updated last year
- Learning Tree structures and Tree metrics☆24Aug 8, 2024Updated 2 years ago
- ☆121Feb 25, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆24Dec 6, 2025Updated 8 months ago
- ☆37Dec 31, 2025Updated 7 months ago
- Code for the paper "Distinguishing the Knowable from the Unknowable with Language Models"☆12Jul 18, 2026Updated last month
- Code for implementing central flows☆49Sep 5, 2025Updated 11 months ago
- Revisiting Character-level Adversarial Attacks for Language Models, ICML 2024☆20Updated this week
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- A high-efficiency text embedding and reranking model based on RWKV architecture.☆20Jan 10, 2026Updated 7 months ago