torch_musa is an open source repository based on PyTorch, which can make full use of the super computing power of MooreThreads graphics cards.
β513Sep 9, 2026Updated last week
Alternatives and similar repositories for torch_musa
Users that are interested in torch_musa are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMsβ121Updated this week
- An adapter layer that ensures torch_musaπ¦ delivers a CUDA-compatible PyTorch experience.β41Updated this week
- MUSA Templates for Linear Algebra Subroutinesβ51Aug 17, 2026Updated last month
- a static analytical model for LLM distributed trainingβ175May 11, 2026Updated 4 months ago
- Get up and running with OpenAI gpt-oss, DeepSeek-R1, Gemma 3 and other models on MTGPU.β37Oct 13, 2025Updated 11 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [HPCA 2026] AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease ofβ¦β385Apr 22, 2026Updated 4 months ago
- a QEMU + gem5 co-simulation framework for AMD MI300X GPU research.β66Sep 11, 2026Updated last week
- Ascend PyTorch adapter (torch_npu). Mirror of https://gitcode.com/Ascend/pytorchβ579Updated this week
- β59Mar 15, 2025Updated last year
- β24Jun 12, 2023Updated 3 years ago
- Examine and discover LoongArch instructionsβ25Sep 5, 2026Updated 2 weeks ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ59Jun 5, 2026Updated 3 months ago
- β23Jan 17, 2026Updated 8 months ago
- Efficient operation implementation based on the Cambricon Machine Learning Unit (MLU) .β182Sep 2, 2026Updated 2 weeks ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- go-onedrive is a Go client library for accessing the Microsoft OneDrive API.β10Dec 12, 2018Updated 7 years ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,431Updated this week
- FlagGems is an operator library for large language models implemented in the Triton Language.β1,121Updated this week
- Development repository for the Triton language and compilerβ20,190Updated this week
- Dissecting NVIDIA GPU Architectureβ128Jul 11, 2022Updated 4 years ago
- GPGPU processor supporting RISCV-V extension, developed with Chisel HDLβ951Aug 18, 2026Updated last month
- β15May 8, 2025Updated last year
- Open Source Computer Vision Libraryβ21Sep 25, 2024Updated last year
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-langβ¦β352Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β15Apr 28, 2026Updated 4 months ago
- chipStar is a tool for compiling and running HIP/CUDA on SPIR-V via OpenCL or Level Zero APIs.β377Updated this week
- study of cutlassβ22Nov 10, 2024Updated last year
- β47Dec 13, 2024Updated last year
- DeepSeek-V3/R1 inference performance simulatorβ195Mar 27, 2025Updated last year
- A model compilation solution for various hardwareβ475Aug 20, 2025Updated last year
- MSCCL++: A GPU-driven communication stack for scalable AI applicationsβ557Updated this week
- Library for modelling performance costs of different Neural Network workloads on NPU devicesβ36Aug 31, 2026Updated 2 weeks ago
- Distributed Compiler and Optimized Parallel Kernelsβ1,546Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.β1,362Aug 28, 2025Updated last year
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,453Updated this week
- The translator that supports translating NVPTX to SPIR-V. This translator is modified from LLVM-SPIR-V Translator.β45Oct 25, 2021Updated 4 years ago
- β11Feb 13, 2025Updated last year
- Xiao's CUDA Optimization Guide [NO LONGER ADDING NEW CONTENT]β333Nov 8, 2022Updated 3 years ago
- Distributed parallel 3D-Causal-VAE for efficient training and inferenceβ50Aug 20, 2025Updated last year
- xserver ddx driver for loongson's display controller and GPUβ11Mar 14, 2023Updated 3 years ago