☆27Jun 8, 2026Updated 2 months ago
Alternatives and similar repositories for torchair
Users that are interested in torchair are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Triton adapter for Ascend. Mirror of https://gitcode.com/ascend/triton-ascend☆127May 18, 2026Updated 2 months ago
- ☆24Jun 10, 2026Updated 2 months ago
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- Ascend PyTorch adapter (torch_npu). Mirror of https://gitcode.com/Ascend/pytorch☆562Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆61Feb 6, 2026Updated 6 months ago
- Simple and efficient memory pool is implemented with C++11.☆10Jun 2, 2022Updated 4 years ago
- ☆92Jan 23, 2025Updated last year
- Sparse kernels for GNNs based on TVM☆17Nov 18, 2020Updated 5 years ago
- Documentation for TCP Lab☆12May 15, 2026Updated 3 months ago
- An Attention Superoptimizer☆22Jan 20, 2025Updated last year
- Ascend TileLang adapter☆350Updated this week
- ☆65Apr 26, 2025Updated last year
- Example of using pytorch's open device registration API☆31Oct 14, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Use contrastive learning to train a large language model (LLM) as a retriever☆12Jul 19, 2024Updated 2 years ago
- ☆19May 9, 2025Updated last year
- ☆18Mar 25, 2026Updated 4 months ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 4 months ago
- Optimized version of dugksFoam with hybrid parallelization strategy and conserved algorithm.☆14Jun 20, 2022Updated 4 years ago
- A GPU algorithm for sparse matrix-matrix multiplication☆74Oct 1, 2020Updated 5 years ago
- ☆37Aug 7, 2025Updated last year
- ☆21Jan 21, 2026Updated 6 months ago
- OpenWrt-6.x for JDC AX1800 Pro | 京东云无线宝亚瑟 AX1800 Pro☆14Jul 29, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- triton for dsa☆69Aug 3, 2026Updated last week
- 使用 cutlass 仓库在 ada 架构上实现 fp8 的 flash attention☆82Aug 12, 2024Updated 2 years ago
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.☆35Mar 18, 2026Updated 4 months ago
- An accelerator to which you can offload RE matching☆14Dec 22, 2024Updated last year
- Multi-Level Triton Runner supporting Python, IR, PTX, AMDGCN, cubin and hasco.☆99May 8, 2026Updated 3 months ago
- ☆12Aug 4, 2022Updated 4 years ago
- [NeurIPS 2024] Efficient LLM Scheduling by Learning to Rank☆82Nov 4, 2024Updated last year
- Composable and Embeddable Communication Runtime for Distributed AI Services☆101Jun 5, 2026Updated 2 months ago
- Transformers components but in Triton☆34May 9, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆150Aug 18, 2025Updated 11 months ago
- [NeurIPS 2022] "NSNet: A General Neural Probabilistic Framework for Satisfiability Problems"☆19Mar 29, 2023Updated 3 years ago
- Ask question to your PDF☆10Jun 11, 2023Updated 3 years ago
- FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation [Efficient ML Model]☆52Aug 2, 2026Updated last week
- Xilinx Modifications to Halide☆13May 3, 2021Updated 5 years ago
- A source-to-source compiler for optimizing CUDA dynamic parallelism by aggregating launches☆15Jun 21, 2019Updated 7 years ago
- AiTer Optimized Model☆156Updated this week