☆26Jun 8, 2026Updated last month
Alternatives and similar repositories for torchair
Users that are interested in torchair are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Triton adapter for Ascend. Mirror of https://gitcode.com/ascend/triton-ascend☆127May 18, 2026Updated 2 months ago
- ☆22Jun 10, 2026Updated last month
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- Ascend PyTorch adapter (torch_npu). Mirror of https://gitcode.com/Ascend/pytorch☆553Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆61Feb 6, 2026Updated 5 months ago
- Simple and efficient memory pool is implemented with C++11.☆10Jun 2, 2022Updated 4 years ago
- ☆91Jan 23, 2025Updated last year
- Documentation for TCP Lab☆12May 15, 2026Updated 2 months ago
- Sparse kernels for GNNs based on TVM☆17Nov 18, 2020Updated 5 years ago
- An Attention Superoptimizer☆22Jan 20, 2025Updated last year
- Ascend TileLang adapter☆338Updated this week
- ☆65Apr 26, 2025Updated last year
- Example of using pytorch's open device registration API☆31Oct 14, 2022Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Use contrastive learning to train a large language model (LLM) as a retriever☆12Jul 19, 2024Updated 2 years ago
- ☆19May 9, 2025Updated last year
- ☆17Mar 25, 2026Updated 4 months ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 3 months ago
- Optimized version of dugksFoam with hybrid parallelization strategy and conserved algorithm.☆14Jun 20, 2022Updated 4 years ago
- A GPU algorithm for sparse matrix-matrix multiplication☆74Oct 1, 2020Updated 5 years ago
- ☆37Aug 7, 2025Updated 11 months ago
- ☆21Jan 21, 2026Updated 6 months ago
- [CAV 2025] PyEuclid: A Versatile Formal Plane Geometry System in Python☆15Jun 27, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- triton for dsa☆68Jul 10, 2026Updated 2 weeks ago
- Learning problem-solving, logic/set, math, physics, economics through functional programming using Haskell☆19Oct 16, 2015Updated 10 years ago
- 使用 cutlass 仓库在 ada 架构上实现 fp8 的 flash attention☆82Aug 12, 2024Updated last year
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.☆34Mar 18, 2026Updated 4 months ago
- An accelerator to which you can offload RE matching☆14Dec 22, 2024Updated last year
- Multi-Level Triton Runner supporting Python, IR, PTX, AMDGCN, cubin and hasco.☆98May 8, 2026Updated 2 months ago
- ☆12Aug 4, 2022Updated 3 years ago
- [NeurIPS 2024] Efficient LLM Scheduling by Learning to Rank☆81Nov 4, 2024Updated last year
- Composable and Embeddable Communication Runtime for Distributed AI Services☆102Jun 5, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Transformers components but in Triton☆34May 9, 2025Updated last year
- ☆149Aug 18, 2025Updated 11 months ago
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆50Nov 27, 2024Updated last year
- [NeurIPS 2022] "NSNet: A General Neural Probabilistic Framework for Satisfiability Problems"☆19Mar 29, 2023Updated 3 years ago
- FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation [Efficient ML Model]☆52Apr 29, 2026Updated 2 months ago
- Ask question to your PDF☆10Jun 11, 2023Updated 3 years ago
- Xilinx Modifications to Halide☆13May 3, 2021Updated 5 years ago