A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm, fused attention, transformer blocks, KV cache management, autoregressive token generation, and end-to-end GPT-2 inference with profiling and benchmarking.
☆43May 17, 2026Updated 2 months ago
Alternatives and similar repositories for gpt2-inference
Users that are interested in gpt2-inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Custom memory allocator in C++ built from scratch using mmap. Allocates a 1MB memory pool upfront and carves blocks from it to keep all a…☆37Apr 29, 2026Updated 3 months ago
- a personal collection of my notes for ml sys☆113Updated this week
- An educational distributed training and inference library for neural nets using local computing☆74Jun 10, 2026Updated last month
- ☆15Mar 11, 2026Updated 4 months ago
- In-process, multi master, distributed database☆21Jun 18, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 5 months ago
- A PyTorch implementation of the GPT-OSS-20B architecture. All components are coded from scratch: RoPE with YaRN, RMSNorm, SwiGLU with cla…☆238Dec 2, 2025Updated 8 months ago
- Daisytuner Optimizing Compiler Collection (docc)☆22Updated this week
- Synchronize the team repository with the services we use☆20Mar 31, 2025Updated last year
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- A programming language with region-based memory management☆32Updated this week
- 强化学习原理 + 强化学习代码实现 + 强化学习框架 + 强化学习论文☆31Updated this week
- Shared Middle-Layer for Triton Compilation☆27Updated this week
- 自动识别文本中的关键词并加粗处理。☆10Oct 30, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- multi-master-paxos with 3 nodes☆14Apr 11, 2022Updated 4 years ago
- An experimental optimizing compiler written by LLM☆28May 23, 2026Updated 2 months ago
- Step by step implementation of a fast softmax kernel in CUDA☆71Jan 6, 2025Updated last year
- Reverse engineering notes. Personal reference only. Everything here is a best-guess reconstruction.☆69Jul 23, 2026Updated last week
- Repository for the CUDA H100 Course☆68Apr 12, 2026Updated 3 months ago
- Andrej Kapathy's micrograd implemented in c☆30Aug 7, 2024Updated last year
- build llm inference code from scratch☆32Apr 21, 2026Updated 3 months ago
- Fast semantic search for biorXiv manuscripts☆12Feb 16, 2025Updated last year
- Building your own docker like container by using only linux bash commands☆16Oct 17, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Tools for visualizing neural nets☆22Jul 29, 2025Updated last year
- This repository contains my coursework and projects completed during the GPU Programming Specialization offered by Johns Hopkins Universi…☆11Jun 13, 2023Updated 3 years ago
- Class to view Alphafold models☆11Jan 20, 2023Updated 3 years ago
- ☆23Jul 6, 2026Updated 3 weeks ago
- real time recommendation playground☆15Nov 7, 2022Updated 3 years ago
- Statistics from our binary transformation framework☆12Jan 16, 2025Updated last year
- ☆12Aug 30, 2024Updated last year
- Compact, local-first neuro-symbolic assistant in Rust with a 200 MiB Bitwork cognitive model, exact reasoning tools, governed memory, and…☆43Jul 24, 2026Updated last week
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Fine-tune FLUX 1.dev for personal AI photos☆22Sep 4, 2024Updated last year
- Official PyTorch implementation for "TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors" [ACL 2026]☆48Apr 14, 2026Updated 3 months ago
- ☆20May 11, 2026Updated 2 months ago
- An imperative command-line-interface for AI workload orchestration☆23Updated this week
- A end-to-end MLOps pipeline for predicting telecom customer churn, featuring automated data preprocessing, ML model training, experiment …☆17Oct 20, 2025Updated 9 months ago
- Splice and merge videos from the terminal☆25Oct 4, 2025Updated 9 months ago
- BatFetch is a command-line tool that displays detailed information about the battery of your device in a clean and organized way.☆24Aug 8, 2024Updated last year