πA curated list of Awesome LLM/VLM Inference Papers with codes: WINT8/4, FlashAttention, PagedAttention, MLA, Parallelism, etc. ππ
β21Mar 30, 2025Updated last year
Alternatives and similar repositories for Awesome-LLM-Inference
Users that are interested in Awesome-LLM-Inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Training codebase for K2-V2β22Dec 17, 2025Updated 9 months ago
- This repository contains the code for the paper "TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back)β¦β15Feb 25, 2026Updated 7 months ago
- Minimal PyTorch implementation of TP, SP, FSDP and sharded-EMAβ32Nov 27, 2025Updated 10 months ago
- Some funny cute/cuteDSL code snippetsβ35Mar 2, 2026Updated 7 months ago
- A class for synchronizing sensor readings to the system clockβ11Oct 25, 2018Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- π200+ Tensor/CUDA Cores Kernels, β‘οΈflash-attn-mma, β‘οΈhgemm with WMMA, MMA and CuTe (98%~100% TFLOPS of cuBLAS/FA2 ππ).β96Apr 26, 2025Updated last year
- Courses_in_ML_DL5β15Nov 6, 2022Updated 3 years ago
- Introducing: Large Scale Capacity Consensus!β14Nov 1, 2021Updated 4 years ago
- Prototypes and experiments for WG Device Management.β17May 21, 2026Updated 4 months ago
- εΎε₯½η¨ηtnn classify demoβ11Mar 24, 2021Updated 5 years ago
- FocusLLM: Scaling LLMβs Context by Parallel Decodingβ45Dec 8, 2024Updated last year
- Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.β48Jun 11, 2025Updated last year
- UnitEval is a benchmarking and evaluation tools for AutoDev Coder.β14Jan 2, 2024Updated 2 years ago
- β15May 13, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS 2025] Official PyTorch implementation for the paper AutoJudge: Judge Decoding Without Manual Annotationβ21Dec 22, 2025Updated 9 months ago
- IntelliJ platform plugin for Wavefront OBJ formatβ15Aug 16, 2026Updated last month
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer (EMNLP 2025)β12Apr 18, 2025Updated last year
- Advanced job scheduling simulatorβ18Nov 6, 2023Updated 2 years ago
- β11Aug 20, 2025Updated last year
- A Chinese-focused PyTorch framework for exploring Attention Residuals in Qwen3-style causal LMs, with baseline, Block AttnRes, Full AttnRβ¦β21May 3, 2026Updated 5 months ago
- β13Jan 22, 2025Updated last year
- EMNLP 2022: Analyzing and Evaluating Faithfulness in Dialogue Summarizationβ13Mar 20, 2025Updated last year
- use yolov3 onnx model to implement object detectionβ10Apr 25, 2019Updated 7 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β68May 19, 2025Updated last year
- 𧬠The adaptive model routing system for exploration and exploitation.β24Jan 4, 2026Updated 9 months ago
- Tracking acceptance rates at global CS/AI conferences. This repository contains metadata for building the OpenAccept main site.β23Updated this week
- ONNX Runtime tiny wrapper for openFrameworksβ15Jan 21, 2022Updated 4 years ago
- β40May 20, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMsβ17Jun 3, 2024Updated 2 years ago
- β141Jun 6, 2025Updated last year
- πA curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.πβ592Jun 13, 2026Updated 3 months ago
- Dockerized container for MODNet - a Real-Time Portrait Matting solutionβ13Mar 27, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- This is the oficial repository for "Safer-Instruct: Aligning Language Models with Automated Preference Data"β17Feb 22, 2024Updated 2 years ago
- β13Jan 7, 2025Updated last year
- β45May 4, 2025Updated last year
- Implemetation of "Pixel-In-Pixel Net: Towards Efficient Facial Landmark Detection in the Wild"β11Jul 6, 2023Updated 3 years ago
- sherpa with mlxβ15Aug 2, 2025Updated last year
- [CVPR 2025] DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Modelsβ88Apr 16, 2026Updated 5 months ago
- PiKV: KV Cache Management System for Mixture of Experts [Efficient ML System]β63Aug 17, 2026Updated last month