Specialized Parallel Linear Algebra, providing distributed GEMM functionality for specific matrix distributions with optional GPU acceleration.
☆32Jun 26, 2024Updated 2 years ago
Alternatives and similar repositories for spla
Users that are interested in spla are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The CSCS ReFrame test suite☆17Updated this week
- Porting meshing tools and solvers that deal with unstructured meshes on GPUs☆15Apr 21, 2026Updated 4 months ago
- ☆24Aug 20, 2026Updated last month
- DLA-Future☆86Jun 19, 2026Updated 3 months ago
- Base container for developing C++ and Fortran HPC applications☆18Jun 14, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Use tensor core to calculate back-to-back HGEMM (half-precision general matrix multiplication) with MMA PTX instruction.☆13Nov 3, 2023Updated 2 years ago
- Runs a single CUDA/OpenCL kernel, taking its source from a file and arguments from the command-line☆26Updated this week
- Domain specific library for electronic structure calculations☆171Updated this week
- MPI+Kokkos implementation of spectral difference method (SDM) high order schemes☆30Feb 2, 2025Updated last year
- STREAMer: Benchmarking remote volatile and non-volatile memory bandwidth☆18Aug 21, 2023Updated 3 years ago
- Netlib Scalapack with robust CMake☆14Mar 26, 2026Updated 5 months ago
- Tensor Algebra for many-body methods☆18Sep 2, 2026Updated 2 weeks ago
- Frame-to-Frame Registration using Gaussian Mixture Models.☆23Mar 2, 2024Updated 2 years ago
- An expression template based linear algebra library running completely on the GPU using CUDA☆26Jun 24, 2021Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A SCVT mesh generation tool☆13Nov 28, 2020Updated 5 years ago
- Subset of BLAS routines optimized for NVIDIA GPUs☆80Mar 27, 2023Updated 3 years ago
- Sparse 3D FFT library with MPI, OpenMP, CUDA and ROCm support☆56Jul 25, 2025Updated last year
- Recipes for software stacks on Alps vClusters.☆16Updated this week
- C++17 Wrapper for ScaLAPACK☆11Oct 5, 2023Updated 2 years ago
- BLAS++ is a C++ wrapper around CPU and GPU BLAS (basic linear algebra subroutines), developed as part of the SLATE project.☆99Aug 27, 2026Updated 3 weeks ago
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.☆33Apr 2, 2025Updated last year
- Generic procedures in Fortran☆22Jan 8, 2019Updated 7 years ago
- Software for Generalized Tensor Decompositions☆19Aug 13, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- HyTeG (Hybrid Tetrahedral Grids) is a C++ framework for large scale high performance finite element simulations based on (but not limited…☆21Jul 2, 2026Updated 2 months ago
- C++ library for graph ordering☆15Mar 20, 2020Updated 6 years ago
- GPU-Accelerated multigrid solver for Poisson's equation in 2D☆31Apr 2, 2026Updated 5 months ago
- Development/testing repo for SWIG+Fortran☆11Mar 25, 2018Updated 8 years ago
- ☆11Aug 8, 2021Updated 5 years ago
- DBCSR: Distributed Block Compressed Sparse Row matrix library☆156Updated this week
- A Monte Carlo Neutron Transport Mini-App☆15Apr 15, 2019Updated 7 years ago
- LAPACK++ is a C++ wrapper around CPU and GPU LAPACK and LAPACK-like linear algebra libraries, developed as part of the SLATE project.☆78Aug 27, 2026Updated 3 weeks ago
- OpenMP offload playground☆10Nov 16, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- libmpdata++ - a library of parallel MPDATA-based solvers for systems of generalised transport equations☆12Jun 19, 2026Updated 3 months ago
- Generic exascale-ready library for halo-exchange operations on variety of grids/meshes☆11Aug 12, 2026Updated last month
- GTensor is a multi-dimensional array C++14 header-only library for hybrid GPU development.☆37Mar 5, 2026Updated 6 months ago
- Massively Asynchronous Coding Environment☆18Oct 21, 2012Updated 13 years ago
- The Kokkos Fortran Interop repository contains tools and interfaces which help interactions between Fortran portions of an applications a…☆47Mar 12, 2026Updated 6 months ago
- 方便扩展的Cuda算子理解和优化框架,仅用在学习使用☆18Jun 13, 2024Updated 2 years ago
- ☆14Sep 22, 2019Updated 6 years ago