common in-memory tensor structure
☆1,232Jun 19, 2026Updated last month
Alternatives and similar repositories for dlpack
Users that are interested in dlpack are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open Machine Learning Compiler Framework☆13,588Updated this week
- A flexible and efficient deep neural network (DNN) compiler that generates high-performance executable from a DNN model description.☆1,002Sep 19, 2024Updated last year
- Dive into Deep Learning Compiler☆649Jun 19, 2022Updated 4 years ago
- The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.☆1,867Updated this week
- TVM integration into PyTorch☆455Jan 15, 2020Updated 6 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A common bricks library for building scalable and portable distributed machine learning.☆878Updated this week
- oneAPI Deep Neural Network Library (oneDNN)☆4,022Updated this week
- CUDA Templates and Python DSLs for High-Performance Linear Algebra☆10,104Updated this week
- Compiler for Neural Network hardware accelerators☆3,321May 11, 2024Updated 2 years ago
- Development repository for the Triton language and compiler☆19,725Updated this week
- A retargetable MLIR-based machine learning compiler and runtime toolkit.☆3,847Updated this week
- [ARCHIVED] Cooperative primitives for CUDA C++. See https://github.com/NVIDIA/cccl☆1,840Oct 9, 2023Updated 2 years ago
- ☆421Feb 24, 2026Updated 4 months ago
- A Python-level JIT compiler designed to make unmodified PyTorch programs faster.☆1,078Apr 17, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Collective communications library with various primitives for multi-machine training.☆1,437Jul 1, 2026Updated 2 weeks ago
- ☆1,650Sep 11, 2018Updated 7 years ago
- FB (Facebook) + GEMM (General Matrix-Matrix Multiplication) - https://code.fb.com/ml-applications/fbgemm/☆1,570Updated this week
- Optimized primitives for collective multi-GPU communication☆4,892Updated this week
- A tensor-aware point-to-point communication primitive for machine learning☆286Dec 17, 2025Updated 7 months ago
- The Tensor Algebra SuperOptimizer for Deep Learning☆743Jan 26, 2023Updated 3 years ago
- A domain specific language to express machine learning workloads.☆1,767Apr 28, 2023Updated 3 years ago
- Symbolic Expression and Statement Module for new DSLs☆207Oct 6, 2020Updated 5 years ago
- a language for fast, portable data-parallel computation☆6,563Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆249Jul 27, 2025Updated 11 months ago
- A list of awesome compiler projects and papers for tensor computation and deep learning.