Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
☆1,073Aug 4, 2026Updated 2 weeks ago
Alternatives and similar repositories for tiny-vllm
Users that are interested in tiny-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A generative pretrained transformer implementation☆94Jun 29, 2026Updated last month
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆113Jun 18, 2026Updated 2 months ago
- learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen☆4,512Updated this week
- Generate hands-on, multi-part technical tutorials on demand, with LLM skills tuned to make content approachable. Then you work through th…☆1,652Aug 3, 2026Updated 3 weeks ago
- a whirlwind tour to deep learning and deep learning systems☆82Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A curated list of best cuda programming books☆962May 19, 2026Updated 3 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆129Apr 17, 2026Updated 4 months ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆1,466Updated this week
- Nano vLLM☆15,100Apr 26, 2026Updated 3 months ago
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆8,746Updated this week
- The best Claude Code that $200 can buy☆269Apr 6, 2026Updated 4 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,532Mar 19, 2026Updated 5 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆1,952Updated this week
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆240Jun 29, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- nCPU: model-native and tensor-optimized CPU research runtimes with organized workloads, tools, and docs☆655Jul 30, 2026Updated 3 weeks ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆4,816May 17, 2026Updated 3 months ago
- Batched square compact-Householder QR factorization.☆15Jul 2, 2026Updated last month
- Shadow any website for offline viewing, with the JavaScript stripped out☆3,291Aug 10, 2026Updated 2 weeks ago
- ☆499May 30, 2026Updated 2 months ago
- ☆3,359May 5, 2026Updated 3 months ago
- Flux 2 image generation model pure C inference☆1,982Feb 13, 2026Updated 6 months ago
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆43May 17, 2026Updated 3 months ago
- Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference…☆688Apr 16, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated last month
- tiny torch, but close to metal☆130Jun 25, 2026Updated last month
- Educational WIP☆73Feb 16, 2026Updated 6 months ago
- High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.☆465Feb 22, 2026Updated 6 months ago
- Learn CUDA with PyTorch☆375Jun 1, 2026Updated 2 months ago
- GPU programming related news and material links☆2,290Jun 15, 2026Updated 2 months ago
- Enlightener, the cutting-edge Retrieval-Augmented Generation (RAG) system that revolutionizes query responses. By combining the power of …☆13Jul 28, 2025Updated last year
- ☆13Apr 10, 2026Updated 4 months ago
- ☆43Aug 16, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Fast per-request isolation for Linux executables with TinyKVM☆72Mar 6, 2026Updated 5 months ago
- CUPTI based GPU profiling library exposing usdt hooks☆37Jun 30, 2026Updated last month
- lowfat - slim your command output. strips noise, saves tokens.☆568Aug 17, 2026Updated last week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆21,674Updated this week
- Tensor library & inference framework for machine learning☆118Oct 3, 2025Updated 10 months ago
- Terraform-style source-of-truth layer for AI agents: HCL specs, LangGraph codegen, and plan/apply/state for hosted agents.☆64Updated this week
- A secure file system for your agents to execute code☆86Updated this week