☆24Aug 26, 2025Updated last year
Alternatives and similar repositories for ThunderMittens
Users that are interested in ThunderMittens are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- JAX support for tvm-ffi abi☆28May 14, 2026Updated 4 months ago
- A Python DSL for Apple Metal GPU compute☆26Jul 23, 2026Updated 2 months ago
- Microbenchmarking hyperparameter tuning for JAX functions.☆24Sep 24, 2026Updated last week
- Compositional Muon release☆25Jun 5, 2026Updated 4 months ago
- ☆25May 23, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official repository for Parallax (Parameterized Local Linear Attention)☆69Jul 30, 2026Updated 2 months ago
- webgpu autograd library☆39May 24, 2025Updated last year
- Implementation of Contrastive Neuron Attribution for behavioral detection and steering.☆40Jun 1, 2026Updated 4 months ago
- Lean Companion to Axler's Linear Algebra Done Right☆28Sep 27, 2026Updated last week
- ☆19Nov 11, 2025Updated 10 months ago
- Repository for GPU related kernels for learning/testing purposes☆21May 27, 2026Updated 4 months ago
- ☆16Sep 24, 2026Updated last week
- Diaphora, the most advanced Free and Open Source program diffing tool. with bninja support☆16May 14, 2026Updated 4 months ago
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 5 months ago
- CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs☆254Updated this week
- PaiNN in jax☆11Jan 14, 2025Updated last year
- ☆29Apr 7, 2025Updated last year
- ☆51Jul 27, 2026Updated 2 months ago
- Code of the book "Getting started with the Julia Programming Language"☆11Jul 7, 2018Updated 8 years ago
- ☆13Jun 2, 2024Updated 2 years ago
- PyTorch implementation of the Flash Spectral Transform Unit.☆23Sep 19, 2024Updated 2 years ago
- Attention variant with per-channel multiplicative decay☆50Jun 3, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- TŒRN is a hands-on standalone open-source sampler and step sequencer built around the Teensy Microcontroller, a 16×16 RGB matrix, with ei…☆19Updated this week
- Tensor Parallelism with JAX + Shard Map☆11Sep 29, 2023Updated 3 years ago
- ☆19Mar 21, 2025Updated last year
- E(n) Equivariant GNN in jax☆14Aug 31, 2023Updated 3 years ago
- Code and data release for the paper "Seeing the Arrow of Time in Large Multimodal Models"☆17Oct 2, 2025Updated last year
- Automatic differentiation for Triton Kernels☆29Aug 12, 2025Updated last year
- trying to make WebGPU a bit easier to use☆19Jan 9, 2024Updated 2 years ago
- ☆14Apr 26, 2018Updated 8 years ago
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Apr 24, 2023Updated 3 years ago
- FLOPS counter for all your GPU benchmarking needs☆13Aug 8, 2024Updated 2 years ago
- ☆20May 30, 2025Updated last year
- code for 'Do Sparse Autoencoders Capture Concept Manifolds?'☆26May 21, 2026Updated 4 months ago
- The official repo for "OpenMoE 2: Sparse Diffusion Language Models".☆58Dec 28, 2025Updated 9 months ago
- Automated High-Performance GPU Kernel Generation☆139Jun 1, 2026Updated 4 months ago
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆26May 5, 2026Updated 5 months ago