new optimizer
☆20Aug 4, 2024Updated last year
Alternatives and similar repositories for grokadamw
Users that are interested in grokadamw are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆137Aug 19, 2024Updated last year
- Data from the paper "Ghostbuster: Detecting Text Ghostwritten by Large Language Models"☆14May 27, 2024Updated 2 years ago
- A basic pure pytorch implementation of flash attention☆17Oct 28, 2024Updated last year
- ☆21Mar 1, 2023Updated 3 years ago
- CUDA implementation of Wavelet KAN.☆17Jun 8, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆16Feb 6, 2024Updated 2 years ago
- a bunch of rubrics I made in different format and structure for llm judge and other use cases☆16Sep 22, 2025Updated 10 months ago
- Knowledge Graph Generator app☆35Apr 18, 2024Updated 2 years ago
- Experimental GPU language with meta-programming☆31Sep 6, 2024Updated last year
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆16Jun 16, 2024Updated 2 years ago
- One File Tensor Libraries☆31Oct 7, 2025Updated 9 months ago
- This repository contains code for the MicroAdam paper.☆21Dec 14, 2024Updated last year
- Pytorch implementation of the PEER block from the paper, Mixture of A Million Experts, by Xu Owen He at Deepmind☆137Nov 1, 2025Updated 8 months ago
- Blazingly fast neighborhood attention☆15Nov 28, 2023Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Quick ADC☆27May 31, 2019Updated 7 years ago
- ☆34May 14, 2025Updated last year
- ☆125Aug 13, 2024Updated last year
- Repository for Sparse Finetuning of LLMs via modified version of the MosaicML llmfoundry☆43Jan 15, 2024Updated 2 years ago
- Code for the paper "QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models".☆277Nov 3, 2023Updated 2 years ago
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated last year
- ☆20Jul 5, 2024Updated 2 years ago
- Code for this paper "HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts via HyperNetwork"☆33Nov 29, 2023Updated 2 years ago
- ☆13May 7, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆33Nov 11, 2024Updated last year
- Extensive time series analysis of chinese PM2.5 content, using models from ARMA and VAR to LSTMs and dynamic time warping clustering☆12Aug 17, 2019Updated 6 years ago
- sync google contacts with information from the dominos data breach <3☆11May 24, 2021Updated 5 years ago
- This is the official repository for the paper "Flora: Low-Rank Adapters Are Secretly Gradient Compressors" in ICML 2024.☆108Jul 1, 2024Updated 2 years ago
- Webcam demo for SKTBrain/DiscoGAN☆12Sep 11, 2019Updated 6 years ago
- Stable diffusion dedicated Hardware with multiple pipelined processor cores☆14Apr 9, 2026Updated 3 months ago
- ☆60Nov 18, 2025Updated 8 months ago
- Reference implementation of models from Nyonic Model Factory☆12May 13, 2024Updated 2 years ago
- Verified Line-Addressed File Editor☆16Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official PyTorch implementation of CD-MOE☆12Mar 18, 2026Updated 4 months ago
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- The Batched API provides a flexible and efficient way to process multiple requests in a batch, with a primary focus on dynamic batching o…☆161Jul 14, 2025Updated last year
- ☆16Dec 9, 2023Updated 2 years ago
- ☆16Apr 2, 2025Updated last year
- ☆33Sep 10, 2024Updated last year
- ☆22Jan 19, 2024Updated 2 years ago