LaurentMazare / gemm-metalLinks

☆14

Alternatives and similar repositories for gemm-metal

Users that are interested in gemm-metal are comparing it to the libraries listed below

Sorting:

LaurentMazare / glim
☆20Updated 9 months ago
google / jaxonnxruntime
A user-friendly tool chain that enables the seamless execution of ONNX models using JAX as the backend.
☆115Updated 3 weeks ago
huggingface / candle-paged-attention
☆12Updated last year
iree-org / iree-jax
☆52Updated 11 months ago
kyutai-labs / jax-flash-attn3
JAX bindings for the flash-attention3 kernels
☆11Updated 11 months ago
kyutai-labs / kaudio
Rust crate for some audio utilities
☆26Updated 4 months ago
LaurentMazare / ug
Experimental compiler for deep learning models
☆68Updated last month
Narsil / bindgen_cuda
☆23Updated 3 months ago
ml-explore / mlx-c
C API for MLX
☆117Updated this week
LaurentMazare / xla-rs
Experimentation using the xla compiler from rust
☆95Updated 11 months ago
cksac / rai
RAI: Rust ML framework with composable transformations like JAX.
☆91Updated 11 months ago
LaurentMazare / tch-ext
Sample Python extension using Rust/PyO3/tch to interact with PyTorch
☆37Updated last year
salykova / sgemm.cu
High-Performance SGEMM on CUDA devices
☆97Updated 5 months ago
kuterd / opal_ptx
Experimental GPU language with meta-programming
☆23Updated 10 months ago
facebookresearch / loop_nest
Loop Nest - Linear algebra compiler and code generator.
☆22Updated 2 years ago
sarah-quinones / gemm
☆89Updated 6 months ago
gevtushenko / llm.c
LLM training in simple, raw C/CUDA
☆99Updated last year
Narsil / ggblas
☆27Updated last year
gpu-mode / discord-cluster-manager
Write a fast kernel and run it on Discord. See how you compare against the best!
☆46Updated this week
moritztng / grayskull-attention
Attention in SRAM on Tenstorrent Grayskull
☆36Updated last year
okuvshynov / llama_duo
asynchronous/distributed speculative evaluation for llama3
☆39Updated 11 months ago
yijunyu / llm.rs
LLM training in simple, raw C/CUDA, migrated into Rust
☆47Updated 3 months ago
EricLBuehler / float8
8-bit floating point types for Rust
☆47Updated 4 months ago
EricLBuehler / candle_graphs
Graph model execution API for Candle
☆13Updated 7 months ago
eugenehp / gpu-fft
GPU based FFT written in Rust and CubeCL
☆23Updated last month
EricLBuehler / zig_ml
Tensor library for Zig
☆11Updated 8 months ago
KerfuffleV2 / smolrsrwkv
A relatively basic implementation of RWKV in Rust written by someone with very little math and ML knowledge. Supports 32, 8 and 4 bit eva…
☆93Updated last year
huggingface / kernel-builder
👷 Build compute kernels
☆77Updated this week
SunDoge / dlpark
A Rust Library for High-Performance Tensor Exchange with Python
☆47Updated last week
FL33TW00D / coremlprofiler
Profile your CoreML models directly from Python 🐍
☆28Updated 9 months ago