Fast matrix multiplication
☆32Jul 6, 2021Updated 5 years ago
Alternatives and similar repositories for fast-matmul
Users that are interested in fast-matmul are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Generating Families of Practical Fast Matrix Multiplication Algorithms☆12Jul 7, 2017Updated 9 years ago
- ☆12Feb 23, 2016Updated 10 years ago
- CNNs in Halide☆22Oct 22, 2015Updated 10 years ago
- ulmBLAS☆111Jun 8, 2025Updated last year
- Artifact of paper "Exploiting Recent SIMD Architectural Advances for Irregular Applications"☆11Jun 23, 2016Updated 10 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Vector Caching Scheme for Streaming FPGA SpMV Accelerators☆10Sep 7, 2015Updated 10 years ago
- Visual graph rewriting platform☆10Jun 3, 2025Updated last year
- ☆11Sep 14, 2020Updated 5 years ago
- BigDataBench Spark workloads☆11Jul 15, 2016Updated 10 years ago
- python module for extracting the system matrices of the fenics fem discretization of (linearized) Navier-Stokes equations and feeding bac…☆10Oct 23, 2025Updated 9 months ago
- Graphs and grammars for Context-Free Path Querying algorithms evaluation.☆11Sep 11, 2024Updated last year
- ☆10Mar 2, 2024Updated 2 years ago
- XMind application packaged into RPM (for Fedora)☆10Dec 9, 2020Updated 5 years ago
- Shared, eMoflon-specific component for incremental unidirectional and bidirectional graph transformations☆17Jun 17, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Nov 1, 2021Updated 4 years ago
- Implementation of the Barnes-Hut algorithm in C++☆12Jul 8, 2011Updated 15 years ago
- ☆16Feb 26, 2018Updated 8 years ago
- chef cookbook to install Apache Spark☆10Jul 17, 2015Updated 11 years ago
- The official website of One Student One Chip project.☆12Feb 5, 2026Updated 6 months ago
- AMDGPU bindings for Flux☆10Apr 6, 2021Updated 5 years ago
- Awesome Geometric Algebra☆31Jul 11, 2020Updated 6 years ago
- ☆12May 3, 2020Updated 6 years ago
- Extended precision arithmetic for Julia (deprecated)☆26Dec 10, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A SystemC + DRAMSim2 simulator for exploring the SpMV hardware accelerator design space.☆15Nov 9, 2014Updated 11 years ago
- Communication-Avoiding Recursive Matrix Multiply☆19Jul 10, 2013Updated 13 years ago
- ☆28Oct 26, 2019Updated 6 years ago
- Heterogeneous Run Time version of MXNet. Added heterogeneous capabilities to the MXNet, uses heterogeneous computing infrastructure frame…☆72Feb 11, 2018Updated 8 years ago
- ☆12Aug 5, 2026Updated last week
- SPMD + Neural Nets☆32Feb 8, 2020Updated 6 years ago
- A Chip Design Automation Solution with Open Source EDA Tools.☆17Updated this week
- ☆18Sep 14, 2021Updated 4 years ago
- OpenDesign Flow Database☆17Oct 31, 2018Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Linux kernel module for triggering a System Management Interrupt (SMI)☆16Sep 13, 2017Updated 8 years ago
- Continuum Dynamics Evaluation and Test Suite☆15Aug 29, 2017Updated 8 years ago
- The fastest tropical matrix multiplication in the world!☆34Dec 1, 2025Updated 8 months ago
- Code-regrouping to reduce latency in Julia code compilation☆16Jun 25, 2026Updated last month
- IMPORTANT NOTICE: This implementation is long outdated. Whole-Function Vectorization is an algorithm that transforms a scalar function in…☆23May 16, 2012Updated 14 years ago
- A CUDA implementation of the Parallel Median of Medians kth element algorithm☆12Jan 29, 2012Updated 14 years ago
- High-Performance Tensor Transpose library☆205May 13, 2023Updated 3 years ago