☆16Sep 24, 2024Updated 2 years ago
Alternatives and similar repositories for py-codegen
Users that are interested in py-codegen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Example of binding a TF32 CUTLASS GEMM kernel to PyTorch☆12Jun 7, 2024Updated 2 years ago
- Programming Gemm Kernels on NVIDIA GPUs with Tensor Cores in Julia☆42Apr 20, 2026Updated 5 months ago
- ☆20May 24, 2025Updated last year
- modified cutlass☆16Oct 26, 2020Updated 5 years ago
- ☆67Feb 24, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A fast implementation of log() and exp()☆58Dec 14, 2022Updated 3 years ago
- SParse AcceleRation on Tensor Architecture☆18Apr 15, 2026Updated 5 months ago
- GPL bases sources for Intel NNP-I card☆18Nov 13, 2023Updated 2 years ago
- ☆13Jun 20, 2019Updated 7 years ago
- study of cutlass☆22Nov 10, 2024Updated last year
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆147Aug 15, 2026Updated last month
- ☆13Feb 5, 2018Updated 8 years ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- Framework for Algorithmic Correctness Testing of Operators☆18Mar 9, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- torch_remat fine-grained activation checkpointing API☆22Updated this week
- A perl script for searching and replacing in mathematics in LaTeX documents.☆13Mar 31, 2026Updated 6 months ago
- Stackfish is an open-source LLM-powered pipeline designed to automatically solve competitive programming problems.☆54Dec 14, 2024Updated last year
- ☆12May 23, 2018Updated 8 years ago
- A Triton-only attention backend for vLLM☆28Jul 14, 2026Updated 2 months ago
- GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs☆16Apr 18, 2025Updated last year
- C-compatible enum for Julia☆14Sep 7, 2026Updated last month
- ☆19Oct 3, 2022Updated 4 years ago
- Code and experiments for the NeurIPS 2023 paper Stabilized Neural Differential Equations for Learning Dynamics with Explicit Constraints☆12Mar 26, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Numerical Analysis I☆23Sep 10, 2026Updated last month
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated 2 years ago
- Support for ternary logic in SSE, XOP, AVX2 and x86 programs☆33Jan 5, 2025Updated last year
- An extension library of WMMA API (Tensor Core API)☆115Jul 12, 2024Updated 2 years ago
- SIMDized check which bytes are in a set☆32Oct 21, 2018Updated 7 years ago
- Tablegen bindings for Rust☆19Sep 25, 2026Updated 2 weeks ago
- Flexible and performant GEMM kernels in Julia☆86Oct 3, 2026Updated last week
- ☆12Sep 25, 2023Updated 3 years ago
- cuASR: CUDA Algebra for Semirings☆52Aug 22, 2022Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Collection of tools/data used for reverse engineering Nintendo Switch sysmodules with Ghidra☆21Jul 10, 2026Updated 3 months ago
- Attention in SRAM on Tenstorrent Grayskull☆38Jul 18, 2024Updated 2 years ago
- CUDA Template Functions☆20Dec 16, 2025Updated 9 months ago
- An arithmetic coder for Rust.☆23May 24, 2023Updated 3 years ago
- ☆84Dec 2, 2022Updated 3 years ago
- A C++ Library for Hardware Design and Simulation☆15May 17, 2020Updated 6 years ago
- 中国人的性格 万泽校对和点评☆15Oct 30, 2014Updated 11 years ago