Accelerated General (FP32) Matrix Multiplication from scratch in CUDA
☆196Jan 9, 2025Updated last year
Alternatives and similar repositories for xGeMM
Users that are interested in xGeMM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- General Matrix Multiplication using NVIDIA Tensor Cores☆29Jan 25, 2025Updated last year
- a simple variational auto encoder with some exploration☆14Nov 22, 2024Updated last year
- A Gentle Introduction to Transformers Neural Network☆15Mar 3, 2024Updated 2 years ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated last year
- Functional Verification of Physical Layer of PCI Express Gen5.0 Graduation Project Using UVM☆35Jul 17, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Code for 'Contrastive Multi-Document Question Generation'☆11Oct 16, 2022Updated 3 years ago
- High-Performance FP32 GEMM on CUDA devices☆128Jan 21, 2025Updated last year
- ☆25Apr 4, 2026Updated 5 months ago
- Super fast FP32 matrix multiplication on RDNA3☆92Mar 30, 2025Updated last year
- Proxy server for triton gRPC server that inferences embedding model in Rust☆21Aug 10, 2024Updated 2 years ago
- AXI Verification using UVM Testbench☆16Oct 20, 2024Updated last year
- A nim module to handle polynomials☆13Jun 7, 2022Updated 4 years ago
- Tensor Algebra for many-body methods☆18Sep 2, 2026Updated 2 weeks ago
- CUTLASS and CuTe Examples☆140Nov 30, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- IO engine for Nim.☆10Jul 8, 2024Updated 2 years ago
- RISC-V vector and tensor compute extensions for Vortex GPGPU acceleration for ML workloads. Optimized for transformer models, CNNs, and g…☆25Apr 25, 2025Updated last year
- Used FPGA board and System Verilog to design controller, DMA, pipelined SIMD processor, and GEMM accelerator☆13Aug 26, 2023Updated 3 years ago
- Repository for GPU related kernels for learning/testing purposes☆21May 27, 2026Updated 3 months ago
- Tutorial introduction slides to GANs. Code implementations and links of relevant papers.☆15May 11, 2019Updated 7 years ago
- fast bpe tokenizer, simple to understand, easy to use☆28Jun 12, 2023Updated 3 years ago
- upcoming concurrent library for Nim☆11Apr 25, 2021Updated 5 years ago
- Nano vLLM☆13Jun 26, 2025Updated last year
- 🧻 Unroll for-loops at compile-time.☆12Jul 27, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- From Minimal GEMM to Everything☆243Jul 9, 2026Updated 2 months ago
- ☆15May 30, 2019Updated 7 years ago
- CUDA Matrix Multiplication Optimization☆281Jul 19, 2024Updated 2 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 7 months ago
- Fastai community entry to 2020 Reproducibility Challenge☆17Oct 20, 2022Updated 3 years ago
- 4k intro sample code written with Nim programming language.☆14Jun 19, 2023Updated 3 years ago
- Flash Attention from Scratch on CUDA Ampere☆197Sep 1, 2025Updated last year
- Learning about CUDA by writing PTX code.☆162Feb 27, 2024Updated 2 years ago
- Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X☆80Feb 11, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Personal solutions to the Triton Puzzles☆22Jul 18, 2024Updated 2 years ago
- ☆20Jan 13, 2026Updated 8 months ago
- Code for ICML2020 "Sequence Generation with Mixed Representations"☆12Jun 27, 2020Updated 6 years ago
- Single page documentation sites☆16May 18, 2021Updated 5 years ago
- 1G‑style Ethernet PHY PCS block in Sky130 using OpenLane☆18Feb 15, 2026Updated 7 months ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- A place to store information for the tensor discussions and possible specifications.☆24Jun 17, 2026Updated 3 months ago