Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
☆1,668Sep 25, 2026Updated this week
Alternatives and similar repositories for AngelSlim
Users that are interested in AngelSlim are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,195Updated this week
- 用手机当鼠标/键盘的极简解决方案☆286May 5, 2026Updated 4 months ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,545Feb 20, 2026Updated 7 months ago
- High Performance LLM Inference Operator Library☆1,171Sep 7, 2026Updated 3 weeks ago
- An independent reimplementation of PowerToys Crop And Lock, featuring Always On Top, rich screenshot annotation, long screenshot, multi-e…☆323Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,130Aug 18, 2026Updated last month
- State-of-the-art LLM compression, built for production inference with vLLM☆3,835Updated this week
- A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包☆1,627Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆858Updated this week
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆5,093Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,525Updated this week
- Cassotis IME - a native Delphi/TSF Chinese Pinyin IME for Windows 10/11. 言泉输入法 —— 基于 Delphi/TSF 的 Windows 10/11 原生开源中文拼音输入法,支持全拼、简拼、六种双拼,…☆267Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,667Updated this week
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A one-stop tool to grab installer packages for VSCode extensions, Chrome/Edge add-ons, Docker images, and Microsoft Store apps — download…☆578Jul 27, 2026Updated 2 months ago
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆7,180Jul 9, 2026Updated 2 months ago
- ☆722Dec 30, 2025Updated 9 months ago
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,265Updated this week
- [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.☆749May 14, 2026Updated 4 months ago
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆3,956Jan 17, 2026Updated 8 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,571Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,862Updated this week
- Model Compression Toolbox for Large Language Models and Diffusion Models☆802Aug 14, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- MicYou is a powerful tool that turns your Android device into a high-quality microphone for your PC.☆4,116Updated this week
- An acceleration library that supports arbitrary bit-width combinatorial quantization operations☆249Sep 30, 2024Updated 2 years ago
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆178Nov 26, 2025Updated 10 months ago
- Global Smooth Auto-Scrolling for Windows / Windows 全局平滑自动滚屏工具(支持多屏协同滚动)☆228Jul 25, 2026Updated 2 months ago
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated 2 months ago
- [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.☆1,257Feb 25, 2026Updated 7 months ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,696Updated this week
- Nano vLLM☆15,673Apr 26, 2026Updated 5 months ago
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,830Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TokenSpeed is a speed-of-light LLM inference engine.☆2,186Updated this week
- ☆826Jun 1, 2026Updated 3 months ago
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Se…☆861Mar 6, 2025Updated last year
- MiniCPM5: SOTA on-device LLMs, small yet powerful.☆11,333Sep 21, 2026Updated last week
- ☆398Apr 16, 2026Updated 5 months ago
- A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.☆1,287Updated this week
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabili…☆4,303Updated this week