in this repository, i'm going to implement increasingly complex llm inference optimizations
☆86May 22, 2025Updated last year
Alternatives and similar repositories for llm-inference-optimizations-explained
Users that are interested in llm-inference-optimizations-explained are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Custom triton kernels for training Karpathy's nanoGPT.☆19Oct 21, 2024Updated last year
- Python module to make coding hassle free!☆10Jun 1, 2021Updated 5 years ago
- aesthetic tensor visualiser☆28Apr 23, 2025Updated last year
- A package for autoencoders/tokenizers also applyable for 3d medical data.☆21Aug 25, 2026Updated 3 weeks ago
- rl from zero pretrain, can it be done? yes.☆296Sep 28, 2025Updated 11 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆15Jan 26, 2025Updated last year
- ~950 line, minimal, extensible LLM inference engine built from scratch.☆482Jan 9, 2026Updated 8 months ago
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- Optimizing Causal LMs through GRPO with weighted reward functions and automated hyperparameter tuning using Optuna☆60Oct 18, 2025Updated 11 months ago
- Retrieve the source code for any model made available on replicate.com!☆36Jan 22, 2024Updated 2 years ago
- A light tensor library in zig.☆77Feb 9, 2025Updated last year
- A curated list of resources for learning and exploring Triton, OpenAI's programming language for writing efficient GPU code.☆498Mar 10, 2025Updated last year
- ☆30Jun 20, 2024Updated 2 years ago
- i will automate factorio☆114Jul 31, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Learning about CUDA by writing PTX code.☆162Feb 27, 2024Updated 2 years ago
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated 2 months ago
- Transformers training in a supercomputer with the 🤗 Stack and Slurm☆15May 9, 2024Updated 2 years ago
- Indoor and outdoor (urban) channel simulations using Winprop API.☆12Dec 2, 2023Updated 2 years ago
- Learnings and programs related to CUDA☆438Jun 29, 2025Updated last year
- Following Karpathy with GPT-2 implementation and training, writing lots of comments cause I have memory of a goldfish☆172Jul 31, 2024Updated 2 years ago
- Collection of autoregressive model implementation☆85Sep 17, 2026Updated last week
- ☆25Oct 10, 2025Updated 11 months ago
- A C++ port of karpathy/micrograd, a tiny scalar-valued autograd engine and a neural net library☆13Nov 24, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- FlexAttention based, minimal vllm-style inference engine for fast Gemma 2 inference.☆360Nov 2, 2025Updated 10 months ago
- Super concise and easily reproducible implementation of a sample LLVM-based compiler with Go.☆19Jan 18, 2026Updated 8 months ago
- ☆81Jun 5, 2024Updated 2 years ago
- Digital Signal Processing LABs for SUSTECH 2020 FALL (EE323).☆15Jan 2, 2021Updated 5 years ago
- In this repository I have a code and brief explanations of the attempts that I made at the ARC-AGI (2024) challenges :)☆26Nov 11, 2024Updated last year
- a tiny vectorstore implementation built with numpy.☆64Apr 26, 2024Updated 2 years ago
- ☆69May 23, 2025Updated last year
- ☆17Apr 29, 2025Updated last year
- 📂 Como um grande fã da Marvel e um apaixonado por tecnologia e jogos, este projeto sem dúvidas é um dos meus favoritos até agora. O proj…☆10Aug 25, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implementation of Fast Weight Attention☆35Sep 17, 2026Updated last week
- So, I trained a Llama a 130M architecture I coded from ground up to build a small instruct model from scratch. Trained on FineWeb dataset…☆17Mar 26, 2025Updated last year
- A tiny easily hackable implementation of a feature dashboard.☆18Oct 21, 2025Updated 11 months ago
- A synthetic story narration dataset to study small audio LMs.☆31Jan 21, 2024Updated 2 years ago
- A Scheduler for Batched LLM Inference☆19Oct 5, 2025Updated 11 months ago
- PTX-Tutorial Written Purely By AIs (Deep Research of Openai and Claude 3.7)☆67Mar 24, 2025Updated last year
- SIMD quantization kernels☆93May 29, 2026Updated 3 months ago