A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm, fused attention, transformer blocks, KV cache management, autoregressive token generation, and end-to-end GPT-2 inference with profiling and benchmarking.
☆44May 17, 2026Updated 4 months ago
Alternatives and similar repositories for gpt2-inference
Users that are interested in gpt2-inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Custom memory allocator in C++ built from scratch using mmap. Allocates a 1MB memory pool upfront and carves blocks from it to keep all a…☆38Apr 29, 2026Updated 5 months ago
- fake CUTLASS to get peformance☆25Apr 28, 2026Updated 5 months ago
- a personal collection of my notes for ml sys☆124Sep 1, 2026Updated last month
- Educational distributed training and inference library for heterogeneous local hardware: Mac minis, Raspberry Pis and GPUs over plain Pyt…☆87Updated this week
- ☆15Mar 11, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Flash Attention from scratch, tiled CUDA forward kernel, online softmax with running max and correction factor, recomputation trick in ba…☆19Mar 6, 2026Updated 7 months ago
- [CoRL'26 Spotlight] Inference engine of "EagleVLA: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference".…☆48Updated this week
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 7 months ago
- A fast image processing TUI written in Rust☆22Aug 18, 2025Updated last year
- Daisytuner Optimizing Compiler Collection (docc)☆23Updated this week
- A static site generator that powers my blog☆19Jun 28, 2025Updated last year
- A curated list of awesome Artificial Life simulators, papers and resources.☆19Dec 9, 2025Updated 9 months ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆28Jan 30, 2026Updated 8 months ago
- Gray-Scott reaction-diffusion system in 3D using CUDA☆12Jun 8, 2019Updated 7 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- This project is a minimal Go implementation of the outbox pattern for an e‑commerce style engine☆22Feb 6, 2026Updated 8 months ago
- A cli tool to download a subfolder of a github repo☆22Jul 19, 2025Updated last year
- 强化学习原理 + 强化学习代码实现 + 强化学习框架 + 强化学习论文☆35Updated this week
- 自动识别文本中的关键词并加粗处理。☆10Oct 30, 2024Updated last year
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆124Feb 28, 2026Updated 7 months ago
- multi-master-paxos with 3 nodes☆14Apr 11, 2022Updated 4 years ago
- Step by step implementation of a fast softmax kernel in CUDA☆71Jan 6, 2025Updated last year
- Andrej Kapathy's micrograd implemented in c☆30Aug 7, 2024Updated 2 years ago
- build llm inference code from scratch☆32Apr 21, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Repository for the CUDA H100 Course☆76Apr 12, 2026Updated 5 months ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- RapidLayout: Fast Hard Block Placement of FPGA-Optimized Systolic Arrays using Evolutionary Algorithms☆20Sep 15, 2026Updated 3 weeks ago
- minimal implementation of kafka paper☆28Apr 7, 2026Updated 6 months ago
- Formalization of the Rupert Problem for convex polyhedra.☆19Sep 10, 2026Updated 3 weeks ago
- ☆24Jul 6, 2026Updated 3 months ago
- Statistics from our binary transformation framework☆12Jan 16, 2025Updated last year
- ☆14Dec 8, 2022Updated 3 years ago
- Conway's Game Of Life implemented using MPI & OpenMP☆13Mar 5, 2019Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A lightweight, modular SQL database engine built from scratch in C.☆35Jul 17, 2026Updated 2 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆83Feb 18, 2026Updated 7 months ago
- Official PyTorch implementation for "TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors" [ACL 2026]☆49Apr 14, 2026Updated 5 months ago
- This repository is a mirror of git://git.kernel.dk/blktrace.git☆15Jun 12, 2016Updated 10 years ago
- ☆45May 4, 2025Updated last year
- 100 days of LLM inference engineering — daily posts, experiments, and visualizations☆1,811Apr 30, 2026Updated 5 months ago
- Inverted triple Pendulum☆18May 13, 2019Updated 7 years ago