Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
☆1,532Aug 7, 2026Updated last week
Alternatives and similar repositories for AngelSlim
Users that are interested in AngelSlim are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,090Updated this week
- 用手机当鼠标/键盘的极简解决方案☆274May 5, 2026Updated 3 months ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,502Feb 20, 2026Updated 5 months ago
- High Performance LLM Inference Operator Library☆1,119Aug 6, 2026Updated last week
- An independent reimplementation of PowerToys Crop And Lock. Always On Top☆319May 12, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,643May 10, 2026Updated 3 months ago
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,696Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆739Updated this week
- A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support…☆1,569Updated this week
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,454Updated this week
- Cassotis IME - a native Delphi/TSF Chinese Pinyin IME for Windows 10/11. 言泉输入法 —— 基于 Delphi/TSF 的 Windows 10/11 原生开源中文拼音输入法,支持全拼、简拼、六种双拼,…☆242Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,187Updated this week
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- SGLang is a high-performance serving framework for large language models and multimodal models.☆32,029Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A one-stop tool to grab installer packages for VSCode extensions, Chrome/Edge add-ons, Docker images, and Microsoft Store apps — download…☆571Jul 27, 2026Updated 3 weeks ago
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆6,993Jul 9, 2026Updated last month
- ☆719Dec 30, 2025Updated 7 months ago
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,234Updated this week
- [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.☆740May 14, 2026Updated 3 months ago
- MicYou is a powerful tool that turns your Android device into a high-quality microphone for your PC.☆3,271Updated this week
- slime is an LLM post-training framework for RL Scaling.☆8,119Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,241Updated this week
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆3,652Jan 17, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Model Compression Toolbox for Large Language Models and Diffusion Models☆799Aug 14, 2025Updated last year
- An acceleration library that supports arbitrary bit-width combinatorial quantization operations☆247Sep 30, 2024Updated last year
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆177Nov 26, 2025Updated 8 months ago
- Global Smooth Auto-Scrolling for Windows / Windows 全局平滑自动滚屏工具(支持多屏协同滚动)☆217Jul 25, 2026Updated 3 weeks ago
- [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.☆1,030Feb 25, 2026Updated 5 months ago
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆132Jul 25, 2026Updated 3 weeks ago
- Nano vLLM☆15,046Apr 26, 2026Updated 3 months ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,313Updated this week
- TokenSpeed is a speed-of-light LLM inference engine.☆1,926Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,708Updated this week
- ☆804Jun 1, 2026Updated 2 months ago
- MiniCPM5-1B: A SOTA 1B on-device LLM, small yet powerful.☆10,197Jul 27, 2026Updated 3 weeks ago
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Se…☆853Mar 6, 2025Updated last year
- ☆394Apr 16, 2026Updated 4 months ago
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabili…☆4,227Updated this week
- A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.☆1,250Updated this week