fmchisel: Efficient Compression and Training Algorithms for Foundation Models
☆90May 4, 2026Updated 3 months ago
Alternatives and similar repositories for fmchisel
Users that are interested in fmchisel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM Daily Summarization of Merged PRs☆53Updated this week
- Cataloging released Triton kernels.☆310Sep 9, 2025Updated 11 months ago
- Pytorch routines for (Ker)nel (Mac)hines☆12Oct 10, 2025Updated 10 months ago
- Make triton easier☆49Jun 12, 2024Updated 2 years ago
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆36Mar 20, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- Efficient Triton Kernels for LLM Training☆6,568Updated this week
- Framework for Algorithmic Correctness Testing of Operators☆16Mar 9, 2026Updated 5 months ago
- A synthetic tabular and relational data generation framework☆72Jul 23, 2026Updated 3 weeks ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆12Nov 8, 2024Updated last year
- vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization☆2,508Updated this week
- ☆10Aug 18, 2016Updated 9 years ago
- Implementation of BitNet-1.58 instruct tuning☆33Apr 14, 2024Updated 2 years ago
- API for coordinating Maintenance in Kubernetes.☆26Jul 2, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- Implementation for FP8/INT8 Rollout for RL training without performence drop.☆309Nov 7, 2025Updated 9 months ago
- ☆24Mar 7, 2025Updated last year
- Transforming Video Diffusion with Temporal Sparse Attention☆56Apr 8, 2026Updated 4 months ago
- Advancing the frontier of efficient AI☆68Jul 10, 2026Updated last month
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆142Sep 24, 2025Updated 10 months ago
- KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches. EMNLP Findings 2024