Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
☆1,138Sep 24, 2026Updated last week
Alternatives and similar repositories for tiny-vllm
Users that are interested in tiny-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A generative pretrained transformer implementation☆94Aug 30, 2026Updated last month
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆123Jun 18, 2026Updated 3 months ago
- Enlightener, the cutting-edge Retrieval-Augmented Generation (RAG) system that revolutionizes query responses. By combining the power of …☆13Jul 28, 2025Updated last year
- ☆13Sep 9, 2026Updated 3 weeks ago
- learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen☆4,746Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Generate hands-on, multi-part technical tutorials on demand, with LLM skills tuned to make content approachable. Then you work through th…☆1,679Aug 3, 2026Updated 2 months ago
- a whirlwind tour to deep learning and deep learning systems☆83Updated this week
- A curated list of best cuda programming books☆985May 19, 2026Updated 4 months ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆3,946Sep 12, 2026Updated 3 weeks ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆252Jun 29, 2026Updated 3 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆131Apr 17, 2026Updated 5 months ago
- Nano vLLM☆15,689Apr 26, 2026Updated 5 months ago
- The best Claude Code that $200 can buy☆268Apr 6, 2026Updated 5 months ago
- Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smar…☆12,945Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆22Sep 24, 2026Updated last week
- TokenSpeed is a speed-of-light LLM inference engine.☆2,189Updated this week
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,577Mar 19, 2026Updated 6 months ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,207May 17, 2026Updated 4 months ago
- nCPU: model-native and tensor-optimized CPU research runtimes with organized workloads, tools, and docs☆660Jul 30, 2026Updated 2 months ago
- Batched square compact-Householder QR factorization.☆15Jul 2, 2026Updated 3 months ago
- ☆577May 30, 2026Updated 4 months ago
- Shadow any website for offline viewing, with the JavaScript stripped out☆3,455Aug 10, 2026Updated last month
- ☆3,418May 5, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆44May 17, 2026Updated 4 months ago
- Flux 2 image generation model pure C inference☆1,993Feb 13, 2026Updated 7 months ago
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated 2 months ago
- Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference…☆689Apr 16, 2026Updated 5 months ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆20Aug 6, 2026Updated last month
- minimal pytorch 4D parallelism☆73Feb 16, 2026Updated 7 months ago
- High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.☆466Feb 22, 2026Updated 7 months ago
- tiny torch, but close to metal☆134Jun 25, 2026Updated 3 months ago
- Learn CUDA with PyTorch☆490Sep 25, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- GPU programming related news and material links☆2,353Jun 15, 2026Updated 3 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆22,863Sep 20, 2026Updated last week
- ☆47Updated this week
- Fast per-request isolation for Linux executables with TinyKVM☆73Mar 6, 2026Updated 6 months ago
- lowfat - slim your command output. strips noise, saves tokens.☆575Sep 24, 2026Updated last week
- Tensor library & inference framework for machine learning☆119Oct 3, 2025Updated last year
- Terraform-style source-of-truth layer for AI agents: HCL specs, LangGraph codegen, and plan/apply/state for hosted agents.☆66Sep 22, 2026Updated last week