learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen
☆4,537Aug 31, 2026Updated this week
Alternatives and similar repositories for tiny-llm
Users that are interested in tiny-llm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Nano vLLM☆15,282Apr 26, 2026Updated 4 months ago
- learn database internals by building a storage engine in Rust☆4,159Aug 14, 2026Updated 2 weeks ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆4,935May 17, 2026Updated 3 months ago
- My learning notes for ML SYS.☆7,160Aug 19, 2026Updated 2 weeks ago
- Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.☆11,874Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- build vector database extensions over Apache Datafusion and/or CMU-DB's BusTub system☆794Updated this week
- An educational OLAP database system.☆1,841Aug 10, 2025Updated last year
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,468Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,318Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆33,326Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆90,787Updated this week
- Learn advanced Rust techniques by building an expression evaluation framework for a database system.☆1,505Updated this week
- slime is an LLM post-training framework for RL Scaling.☆8,354Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,328Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Datacenter Scale Distributed Inference Serving Framework☆7,950Updated this week
- Materials for learning SGLang☆884Jan 5, 2026Updated 7 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,619Updated this week
- Material for gpu-mode lectures☆6,535Jun 15, 2026Updated 2 months ago
- Puzzles for learning Triton, play it with minimal environment configuration!☆746Mar 17, 2026Updated 5 months ago
- Machine Learning Engineering Open Book☆18,880Updated this week
- The best ChatGPT that $100 can buy.☆57,738Aug 2, 2026Updated last month
- The BusTub Relational Database Management System (Educational)☆5,072Updated this week
- 分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等☆3,776Aug 7, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A course to build distributed key-value service based on TiKV model☆3,986May 3, 2025Updated last year
- A high-performance distributed file system designed to address the challenges of AI training and inference workloads.☆10,179May 7, 2026Updated 3 months ago
- 🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for Machine Learning, LLM (Large Language Mod…☆4,330Jul 25, 2025Updated last year
- AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。☆8,093Dec 22, 2025Updated 8 months ago
- Distributed SQL database in Rust, written as an educational project☆7,276Jul 27, 2026Updated last month
- Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O☆596Sep 13, 2025Updated 11 months ago
- MLX: An array framework for Apple silicon☆28,271Updated this week
- Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation☆8,061May 15, 2025Updated last year
- Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2☆669Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- LLM training in simple, raw C/CUDA☆30,914Jun 26, 2025Updated last year
- CUDA Templates and Python DSLs for High-Performance Linear Algebra☆10,362Updated this week
- Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation☆1,010Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,258Updated this week
- 🧠 Train a 64M-parameter LLM from scratch in just 2h!☆57,775Updated this week
- Implement a ChatGPT-like LLM in PyTorch from scratch, step by step☆104,208Updated this week
- Databend 内幕大揭秘☆301Jan 26, 2024Updated 2 years ago