Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
☆1,630Sep 4, 2026Updated this week
Alternatives and similar repositories for AngelSlim
Users that are interested in AngelSlim are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,152Updated this week
- 用手机当鼠标/键盘的极简解决方案☆280May 5, 2026Updated 4 months ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,525Feb 20, 2026Updated 6 months ago
- High Performance LLM Inference Operator Library☆1,146Updated this week
- An independent reimplementation of PowerToys Crop And Lock. Always On Top with rich screenshot annotation, long screenshot, multi-engine …☆324Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,056Aug 18, 2026Updated 2 weeks ago
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,764Updated this week
- A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包☆1,606Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆812Updated this week
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,761Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,343Updated this week
- Cassotis IME - a native Delphi/TSF Chinese Pinyin IME for Windows 10/11. 言泉输入法 —— 基于 Delphi/TSF 的 Windows 10/11 原生开源中文拼音输入法,支持全拼、简拼、六种双拼,…☆256Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆35,594Updated this week
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A one-stop tool to grab installer packages for VSCode extensions, Chrome/Edge add-ons, Docker images, and Microsoft Store apps — download…☆576Jul 27, 2026Updated last month
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆7,084Jul 9, 2026Updated last month
- ☆722Dec 30, 2025Updated 8 months ago
- [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.☆747May 14, 2026Updated 3 months ago
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,252Updated this week
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆3,708Jan 17, 2026Updated 7 months ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,359Updated this week
- slime is an LLM post-training framework for RL Scaling.☆8,404Updated this week
- MicYou is a powerful tool that turns your Android device into a high-quality microphone for your PC.☆3,641Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Model Compression Toolbox for Large Language Models and Diffusion Models☆802Aug 14, 2025Updated last year
- An acceleration library that supports arbitrary bit-width combinatorial quantization operations☆248Sep 30, 2024Updated last year
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆177Nov 26, 2025Updated 9 months ago
- Global Smooth Auto-Scrolling for Windows / Windows 全局平滑自动滚屏工具(支持多屏协同滚动)☆224Jul 25, 2026Updated last month
- [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.☆1,043Feb 25, 2026Updated 6 months ago
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated last month
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,516Updated this week
- Nano vLLM☆15,333Apr 26, 2026Updated 4 months ago
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,756Aug 13, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TokenSpeed is a speed-of-light LLM inference engine.☆2,103Updated this week
- ☆816Jun 1, 2026Updated 3 months ago
- MiniCPM5: SOTA on-device LLMs, small yet powerful.☆10,318Updated this week
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Se…☆856Mar 6, 2025Updated last year
- ☆397Apr 16, 2026Updated 4 months ago
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabili…☆4,275Updated this week
- A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.☆1,273Updated this week