End to End steps for adding custom ops in PyTorch.
☆24Aug 20, 2020Updated 5 years ago
Alternatives and similar repositories for pytorch_custom_op
Users that are interested in pytorch_custom_op are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CAKE Library for constant-bandwidth matrix multiplication on CPUs☆14Apr 6, 2024Updated 2 years ago
- A intelligent matrix format designer for SpMV☆10Oct 10, 2023Updated 2 years ago
- A pseudo random number generator library written against the SYCL API.☆11Jun 11, 2019Updated 7 years ago
- Marek's approach to building AMD GPU drivers for driver development☆28Jun 1, 2026Updated last month
- HierarchicalKV is a part of NVIDIA Merlin and provides hierarchical key-value storage to meet RecSys requirements. The key capability of…☆208May 22, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- My tests and experiments with some popular dl frameworks.☆17Sep 11, 2025Updated 10 months ago
- Example to build PyTorch CUDA extension using CMake (with pybind11 and scikit-build)☆12May 26, 2020Updated 6 years ago
- A simple shellscript for splitting the PDF of a paper into the main body and an appendix.☆18Jun 1, 2020Updated 6 years ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- An Easy To Use PyTorch Computer Vision Library☆52Jul 6, 2023Updated 3 years ago
- ☆23Feb 16, 2022Updated 4 years ago
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- Source code for the paper "LongGenBench: Long-context Generation Benchmark"☆24Oct 8, 2024Updated last year
- Hardware go brrr bounded context suffix array construction algorithm☆19Nov 1, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- High-speed Bloom filters and taffy filters for C, C++, and Java☆35Aug 9, 2023Updated 2 years ago
- TiledKernel is a code generation library based on macro kernels and memory hierarchy graph data structure.☆19May 12, 2024Updated 2 years ago
- Noisy language compiler☆17Jul 31, 2024Updated last year
- Highly Efficient FFT for Exascale☆38Apr 29, 2024Updated 2 years ago
- A study for a custom convolution layer in which the x and y components of an image pixel are added to the kernel inputs.☆12Feb 21, 2020Updated 6 years ago
- Commands that will make you more comfortable with the ROCm toolkit.☆18Aug 1, 2024Updated last year
- Automatic virtualization of (general) accelerators.☆47Nov 28, 2022Updated 3 years ago
- Examples for MS-AMP package.☆30Jul 17, 2025Updated last year
- Yinghan's Code Sample☆365Jul 25, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- TVMScript kernel for deformable attention☆25Dec 15, 2021Updated 4 years ago
- Linux io_uring based c++ 20 coroutine library☆28Jun 21, 2022Updated 4 years ago
- ☆17Updated this week
- Setting up Vscode to work with Pytorch in C/C++ with CUDA support☆25Feb 5, 2025Updated last year
- UCSD CSE 237D Spring '20 Course Project☆21Sep 4, 2023Updated 2 years ago
- DeeperGEMM: crazy optimized version☆86May 5, 2025Updated last year
- ☆33Jul 19, 2024Updated 2 years ago
- PTX-Tutorial Written Purely By AIs (Deep Research of Openai and Claude 3.7)☆66Mar 24, 2025Updated last year
- ☆23Feb 18, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A Simple but Powerful CNN Trainer For PyTorch☆26Dec 2, 2020Updated 5 years ago
- CUDA实现'huawei-noah/AdderNet'的forward和backward☆17Apr 16, 2020Updated 6 years ago
- Implementation of RegNetY in TensorFlow 2☆21Jan 15, 2023Updated 3 years ago
- A toy eval suite for tracing generalization dynamics of LM pre-training☆19May 19, 2026Updated 2 months ago
- Experimental OpenCL SPIR-V to OpenCL C translator☆29Mar 1, 2026Updated 4 months ago
- This is a BNN_Kernel on PyTorch for 1-bit networks in image data processing☆23Sep 28, 2019Updated 6 years ago
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆30Jan 22, 2026Updated 6 months ago