Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
☆963Jul 2, 2026Updated 3 weeks ago
Alternatives and similar repositories for tiny-vllm
Users that are interested in tiny-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A generative pretrained transformer implementation☆93Jun 29, 2026Updated last month
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆108Jun 18, 2026Updated last month
- A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.☆4,425Updated this week
- Generate hands-on, multi-part technical tutorials on demand, with LLM skills tuned to make content approachable. Then you work through th…☆1,599Jul 20, 2026Updated last week
- a whirlwind tour to deep learning and deep learning systems☆81Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A curated list of best cuda programming books☆941May 19, 2026Updated 2 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆123Apr 17, 2026Updated 3 months ago
- A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do☆1,273Apr 27, 2026Updated 3 months ago
- Nano vLLM☆14,679Apr 26, 2026Updated 3 months ago
- Foundation model for tiny devices; 14mb, 26m params, 1-6k toks/sec on mobiles, wearables smart home and robots.☆3,296Updated this week
- The best Claude Code that $200 can buy☆269Apr 6, 2026Updated 3 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,488Mar 19, 2026Updated 4 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆1,751Updated this week
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆234Jun 29, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- nCPU: model-native and tensor-optimized CPU research runtimes with organized workloads, tools, and docs☆655Jul 11, 2026Updated 2 weeks ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆4,645May 17, 2026Updated 2 months ago
- Batched square compact-Householder QR factorization.☆14Jul 2, 2026Updated 3 weeks ago
- Shadow any website for offline viewing, with the JavaScript stripped out☆2,967Jul 11, 2026Updated 2 weeks ago
- ☆494May 30, 2026Updated last month
- ☆3,216May 5, 2026Updated 2 months ago
- Flux 2 image generation model pure C inference☆1,967Feb 13, 2026Updated 5 months ago
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆43May 17, 2026Updated 2 months ago
- Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference…☆688Apr 16, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated 3 weeks ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆17Updated this week
- tiny torch, but close to metal☆130Jun 25, 2026Updated last month
- Educational WIP☆73Feb 16, 2026Updated 5 months ago
- High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.☆464Feb 22, 2026Updated 5 months ago
- Learn CUDA with PyTorch☆360Jun 1, 2026Updated last month
- GPU programming related news and material links☆2,247Jun 15, 2026Updated last month
- Enlightener, the cutting-edge Retrieval-Augmented Generation (RAG) system that revolutionizes query responses. By combining the power of …☆13Jul 28, 2025Updated last year
- ☆13Apr 10, 2026Updated 3 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆35Jun 28, 2026Updated last month
- Fast per-request isolation for Linux executables with TinyKVM☆71Mar 6, 2026Updated 4 months ago
- CUPTI based GPU profiling library exposing usdt hooks☆37Jun 30, 2026Updated 3 weeks ago
- lowfat - slim your command output. strips noise, saves tokens.☆563Jul 8, 2026Updated 3 weeks ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆19,409Updated this week
- Tensor library & inference framework for machine learning☆118Oct 3, 2025Updated 9 months ago
- Terraform-style source-of-truth layer for AI agents: HCL specs, LangGraph codegen, and plan/apply/state for hosted agents.☆60Updated this week