Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O
☆601Sep 13, 2025Updated last year
Alternatives and similar repositories for yalm
Users that are interested in yalm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CUDA/Metal accelerated language model inference☆646May 29, 2025Updated last year
- CPU inference for the DeepSeek family of large language models in C++☆325Oct 2, 2025Updated 11 months ago
- Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.☆12,010Updated this week
- Examples of CUDA implementations by Cutlass CuTe☆283Jul 1, 2025Updated last year
- 使用 CUDA C++ 实 现的 llama 模型推理框架☆63Nov 8, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Nano vLLM☆15,636Apr 26, 2026Updated 5 months ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,182May 17, 2026Updated 4 months ago
- Static suckless single batch CUDA-only qwen3-0.6B mini inference engine☆560Sep 8, 2025Updated last year
- flash attention tutorial written in python, triton, cuda, cutlass☆535Jan 20, 2026Updated 8 months ago
- FlashInfer: Kernel Library for LLM Serving☆6,510Updated this week
- ☆192May 11, 2026Updated 4 months ago
- Flash Attention in ~100 lines of CUDA (forward pass only)☆1,190Dec 30, 2024Updated last year
- Material for gpu-mode lectures☆6,658Sep 9, 2026Updated 2 weeks ago
- Fast low-bit matmul kernels in Triton☆487Aug 9, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Fast CUDA matrix multiplication from scratch☆1,320Sep 2, 2025Updated last year
- ⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.☆158May 10, 2025Updated last year
- Kernels, of the mega variety :)☆835May 26, 2026Updated 4 months ago
- 📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉☆5,515Aug 14, 2026Updated last month
- Sparse & quantized LLM training/inference/CPT/SFT/DPO☆28Sep 13, 2026Updated 2 weeks ago
- CPM.cu is a lightweight, high-performance CUDA implementation for LLMs, optimized for end-device inference and featuring cutting-edge tec…☆243Jan 14, 2026Updated 8 months ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated last year
- High Performance LLM Inference Operator Library☆1,168Sep 7, 2026Updated 3 weeks ago
- Tile primitives for speedy kernels☆3,730Sep 12, 2026Updated 2 weeks ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Implement Flash Attention using Cute.