Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
☆1,491Jul 29, 2026Updated this week
Alternatives and similar repositories for AngelSlim
Users that are interested in AngelSlim are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,027Updated this week
- 用手机当鼠标/键盘的极简解决方案☆268May 5, 2026Updated 2 months ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,481Feb 20, 2026Updated 5 months ago
- High Performance LLM Inference Operator Library☆1,070Updated this week
- An independent reimplementation of PowerToys Crop And Lock. Always On Top☆316May 12, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,547May 10, 2026Updated 2 months ago
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,602Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆669Updated this week
- A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support…☆1,543Updated this week
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,341Updated this week
- Cassotis IME - a native Delphi/TSF Chinese Pinyin IME for Windows 10/11. 言泉输入法 —— 基于 Delphi/TSF 的 Windows 10/11 原生开源中文拼音输入法,支持全拼、简拼、六种双拼,…☆232Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,062Updated this week
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated 11 months ago
- SGLang is a high-performance serving framework for large language models and multimodal models.☆30,926Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A one-stop tool to grab installer packages for VSCode extensions, Chrome/Edge add-ons, Docker images, and Microsoft Store apps — download…☆559Updated this week
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆6,810Jul 9, 2026Updated 3 weeks ago
- ☆713Dec 30, 2025Updated 6 months ago
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,217Updated this week
- [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.☆736May 14, 2026Updated 2 months ago
- MicYou is a powerful tool that turns your Android device into a high-quality microphone for your PC.☆3,131Updated this week
- slime is an LLM post-training framework for RL Scaling.☆7,696Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,007Updated this week
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆3,518Jan 17, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Model Compression Toolbox for Large Language Models and Diffusion Models☆796Aug 14, 2025Updated 11 months ago
- An acceleration library that supports arbitrary bit-width combinatorial quantization operations☆246Sep 30, 2024Updated last year
- Global Smooth Auto-Scrolling for Windows / Windows 全局平滑自动滚屏工具(支持多屏协同滚动)☆214Updated this week
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆176Nov 26, 2025Updated 8 months ago
- [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.☆1,019Feb 25, 2026Updated 5 months ago
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆124Updated this week
- Nano vLLM☆14,679Apr 26, 2026Updated 3 months ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,083Updated this week
- TokenSpeed is a speed-of-light LLM inference engine.☆1,751Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,600Jul 14, 2026Updated 2 weeks ago
- ☆797Jun 1, 2026Updated last month
- MiniCPM5-1B: A SOTA 1B on-device LLM, small yet powerful.☆10,064Updated this week
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Se…☆852Mar 6, 2025Updated last year
- ☆390Apr 16, 2026Updated 3 months ago
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabili…☆4,198Updated this week
- A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.☆1,239Updated this week