new optimizer
☆20Aug 4, 2024Updated 2 years ago
Alternatives and similar repositories for grokadamw
Users that are interested in grokadamw are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆137Aug 19, 2024Updated 2 years ago
- Data from the paper "Ghostbuster: Detecting Text Ghostwritten by Large Language Models"☆14May 27, 2024Updated 2 years ago
- A basic pure pytorch implementation of flash attention☆17Oct 28, 2024Updated last year
- ☆21Mar 1, 2023Updated 3 years ago
- CUDA implementation of Wavelet KAN.☆17Jun 8, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆16Feb 6, 2024Updated 2 years ago
- Experimental GPU language with meta-programming☆32Sep 6, 2024Updated last year
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆16Jun 16, 2024Updated 2 years ago
- One File Tensor Libraries☆31Oct 7, 2025Updated 10 months ago
- A byte-level decoder architecture that matches the performance of tokenized Transformers.☆68Apr 24, 2024Updated 2 years ago
- ☆10Jul 8, 2019Updated 7 years ago
- This repository contains code for the MicroAdam paper.☆21Dec 14, 2024Updated last year
- Pytorch implementation of the PEER block from the paper, Mixture of A Million Experts, by Xu Owen He at Deepmind☆138Nov 1, 2025Updated 9 months ago
- Blazingly fast neighborhood attention☆15Nov 28, 2023Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆14Dec 11, 2022Updated 3 years ago
- ☆34May 14, 2025Updated last year
- ☆125Aug 13, 2024Updated 2 years ago
- Repository for Sparse Finetuning of LLMs via modified version of the MosaicML llmfoundry☆43Jan 15, 2024Updated 2 years ago
- Code for the paper "QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models".☆277Nov 3, 2023Updated 2 years ago
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated 2 years ago
- Code for this paper "HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts via HyperNetwork"☆33Nov 29, 2023Updated 2 years ago
- ☆13May 7, 2023Updated 3 years ago
- ☆33Nov 11, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Extensive time series analysis of chinese PM2.5 content, using models from ARMA and VAR to LSTMs and dynamic time warping clustering☆12Aug 17, 2019Updated 7 years ago
- sync google contacts with information from the dominos data breach <3☆11May 24, 2021Updated 5 years ago
- Fast, High-Fidelity LLM Decoding with Regex Constraints☆21Jul 26, 2024Updated 2 years ago
- This is the official repository for the paper "Flora: Low-Rank Adapters Are Secretly Gradient Compressors" in ICML 2024.☆108Jul 1, 2024Updated 2 years ago
- ☆60Nov 18, 2025Updated 9 months ago
- Verified Line-Addressed File Editor☆16Updated this week
- GGML bindings that aim to be idiomatic Rust rather than directly corresponding to the C/C++ interface☆20Sep 25, 2023Updated 2 years ago
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- Give AI agents ambitious work without losing the plot. nac is an open-source harness for long-running tasks, using a central orchestrator…☆140Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Clustered Compositional Embeddings☆13Oct 25, 2023Updated 2 years ago
- Analyzing LLM Alignment via Token distribution shift☆17Jan 26, 2024Updated 2 years ago
- The Batched API provides a flexible and efficient way to process multiple requests in a batch, with a primary focus on dynamic batching o…☆162Jul 14, 2025Updated last year
- ☆16Dec 9, 2023Updated 2 years ago
- ☆16Apr 2, 2025Updated last year
- ☆33Sep 10, 2024Updated last year
- Official repo: “Oh LLM, I’m Asking Thee, Please Give Me a Decision Tree”: Zero-Shot Decision Tree Induction and Embedding with Large Lang…☆17Aug 10, 2026Updated last week