torch_musa is an open source repository based on PyTorch, which can make full use of the super computing power of MooreThreads graphics cards.
β508Aug 17, 2026Updated last week
Alternatives and similar repositories for torch_musa
Users that are interested in torch_musa are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMsβ116Updated this week
- An adapter layer that ensures torch_musaπ¦ delivers a CUDA-compatible PyTorch experience.β40Updated this week
- MUSA Templates for Linear Algebra Subroutinesβ49Aug 17, 2026Updated last week
- a static analytical model for LLM distributed trainingβ174May 11, 2026Updated 3 months ago
- [HPCA 2026] AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease ofβ¦β377Apr 22, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- a QEMU + gem5 co-simulation framework for AMD MI300X GPU research.β61Aug 15, 2026Updated 2 weeks ago
- Ascend PyTorch adapter (torch_npu). Mirror of https://gitcode.com/Ascend/pytorchβ567Updated this week
- β57Mar 15, 2025Updated last year
- β24Jun 12, 2023Updated 3 years ago
- Examine and discover LoongArch instructionsβ25Jun 4, 2026Updated 2 months ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ58Jun 5, 2026Updated 2 months ago
- β23Jan 17, 2026Updated 7 months ago
- Efficient operation implementation based on the Cambricon Machine Learning Unit (MLU) .β180Updated this week
- go-onedrive is a Go client library for accessing the Microsoft OneDrive API.β10Dec 12, 2018Updated 7 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β13Jan 3, 2025Updated last year
- β14May 8, 2025Updated last year
- Triton for OpenCL backend, and use mlir-translate to get source OpenCL codeβ27Aug 27, 2025Updated last year
- Dissecting NVIDIA GPU Architectureβ127Jul 11, 2022Updated 4 years ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,302Updated this week
- FlagGems is an operator library for large language models implemented in the Triton Language.β1,086Updated this week
- Development repository for the Triton language and compilerβ20,037Updated this week
- GPGPU processor supporting RISCV-V extension, developed with Chisel HDLβ940Aug 18, 2026Updated last week
- An unofficial cuda assembler, for all generations of SASS, hopefully οΌοΌβ621Apr 20, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Open Source Computer Vision Libraryβ21Sep 25, 2024Updated last year
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-langβ¦β322Updated this week
- β15Apr 28, 2026Updated 4 months ago
- chipStar is a tool for compiling and running HIP/CUDA on SPIR-V via OpenCL or Level Zero APIs.β372Updated this week
- study of cutlassβ22Nov 10, 2024Updated last year
- β47Dec 13, 2024Updated last year
- DeepSeek-V3/R1 inference performance simulatorβ195Mar 27, 2025Updated last year
- A model compilation solution for various hardwareβ475Aug 20, 2025Updated last year
- MSCCL++: A GPU-driven communication stack for scalable AI applicationsβ550Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Library for modelling performance costs of different Neural Network workloads on NPU devicesβ36Jul 23, 2026Updated last month
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.β1,355Aug 28, 2025Updated last year
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,345Updated this week
- The translator that supports translating NVPTX to SPIR-V. This translator is modified from LLVM-SPIR-V Translator.β45Oct 25, 2021Updated 4 years ago
- β11Feb 13, 2025Updated last year
- Xiao's CUDA Optimization Guide [NO LONGER ADDING NEW CONTENT]β333Nov 8, 2022Updated 3 years ago
- xserver ddx driver for loongson's display controller and GPUβ11Mar 14, 2023Updated 3 years ago