Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
☆1,106Aug 23, 2026Updated 3 weeks ago
Alternatives and similar repositories for tiny-vllm
Users that are interested in tiny-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A generative pretrained transformer implementation☆94Aug 30, 2026Updated 2 weeks ago
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆115Jun 18, 2026Updated 2 months ago
- Enlightener, the cutting-edge Retrieval-Augmented Generation (RAG) system that revolutionizes query responses. By combining the power of …☆13Jul 28, 2025Updated last year
- ☆13Updated this week
- learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen☆4,560Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Generate hands-on, multi-part technical tutorials on demand, with LLM skills tuned to make content approachable. Then you work through th…☆1,664Aug 3, 2026Updated last month
- a whirlwind tour to deep learning and deep learning systems☆83Sep 2, 2026Updated last week
- A curated list of best cuda programming books☆977May 19, 2026Updated 3 months ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆2,741Aug 23, 2026Updated 3 weeks ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆242Jun 29, 2026Updated 2 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆129Apr 17, 2026Updated 4 months ago
- Nano vLLM☆15,411Apr 26, 2026Updated 4 months ago
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆10,888Updated this week
- The best Claude Code that $200 can buy☆269Apr 6, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,555Mar 19, 2026Updated 5 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆2,116Updated this week
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,032May 17, 2026Updated 3 months ago
- nCPU: model-native and tensor-optimized CPU research runtimes with organized workloads, tools, and docs☆657Jul 30, 2026Updated last month
- Batched square compact-Householder QR factorization.☆15Jul 2, 2026Updated 2 months ago
- ☆576May 30, 2026Updated 3 months ago
- Shadow any website for offline viewing, with the JavaScript stripped out☆3,398Aug 10, 2026Updated last month
- ☆3,405May 5, 2026Updated 4 months ago
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆44May 17, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Flux 2 image generation model pure C inference☆1,991Feb 13, 2026Updated 6 months ago
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated 2 months ago
- Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference…☆688Apr 16, 2026Updated 4 months ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆20Aug 6, 2026Updated last month
- minimal pytorch 4D parallelism☆73Feb 16, 2026Updated 6 months ago
- High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.☆465Feb 22, 2026Updated 6 months ago
- tiny torch, but close to metal☆132Jun 25, 2026Updated 2 months ago
- Learn CUDA with PyTorch☆469Updated this week
- GPU programming related news and material links☆2,329Jun 15, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆44Aug 16, 2026Updated 3 weeks ago
- Fast per-request isolation for Linux executables with TinyKVM☆73Mar 6, 2026Updated 6 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆22,297Updated this week
- CUPTI based GPU profiling library exposing usdt hooks☆38Jun 30, 2026Updated 2 months ago
- lowfat - slim your command output. strips noise, saves tokens.☆571Aug 17, 2026Updated 3 weeks ago
- Tensor library & inference framework for machine learning☆118Oct 3, 2025Updated 11 months ago
- Terraform-style source-of-truth layer for AI agents: HCL specs, LangGraph codegen, and plan/apply/state for hosted agents.☆65Updated this week