Small scale distributed training of sequential deep learning models, built on Numpy and MPI.
☆165Oct 19, 2023Updated 2 years ago
Alternatives and similar repositories for ShallowSpeed
Users that are interested in ShallowSpeed are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆93Jul 5, 2024Updated 2 years ago
- Fast CUDA matrix multiplication from scratch☆1,256Sep 2, 2025Updated 10 months ago
- Step-by-step optimization of CUDA SGEMM☆485Mar 30, 2022Updated 4 years ago
- Minimal but scalable implementation of large language models in JAX☆34Nov 28, 2025Updated 7 months ago
- a whirlwind tour to deep learning and deep learning systems☆81Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆591Jul 11, 2024Updated 2 years ago
- Fastest kernels written from scratch☆583Sep 18, 2025Updated 10 months ago
- What would you do with 1000 H100s...☆1,181Jan 10, 2024Updated 2 years ago
- 🔀 yet another mixture of experts☆23Jun 5, 2026Updated last month
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,254Aug 26, 2025Updated 10 months ago
- GPU programming related news and material links☆2,233Jun 15, 2026Updated last month
- a tiny vectorstore implementation built with numpy.☆64Apr 26, 2024Updated 2 years ago
- ☆60Feb 24, 2026Updated 4 months ago
- ☆30Dec 2, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆46May 24, 2025Updated last year
- writing really fast kernels☆19Updated this week
- Algorithms for approximate attention in LLMs☆22Apr 14, 2025Updated last year
- ☆170Dec 27, 2024Updated last year
- Lighter than feather, Harder than steel, RL Framework☆24Mar 26, 2026Updated 3 months ago
- coloring terminal text with intensities (used for plotting probability, entropy with tokens)☆12Oct 11, 2024Updated last year
- Puzzles for learning Triton☆2,531Apr 1, 2026Updated 3 months ago
- ring-attention experiments☆171Oct 17, 2024Updated last year
- A repository to unravel the language of GPUs, making their kernel conversations easy to understand☆208Jun 1, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Various reinforcement learning algorithms written in Jax + Flax☆26Jun 24, 2023Updated 3 years ago
- UNet diffusion model in pure CUDA☆661Jun 28, 2024Updated 2 years ago
- Large scale 4D parallelism pre-training for 🤗 transformers in Mixture of Experts *(still work in progress)*☆87Dec 14, 2023Updated 2 years ago
- Cataloging released Triton kernels.☆311Sep 9, 2025Updated 10 months ago
- ☆51Jan 18, 2024Updated 2 years ago
- Tile primitives for speedy kernels☆3,552Jul 13, 2026Updated last week
- run paligemma in real time☆132May 18, 2024Updated 2 years ago
- ☆22Jan 22, 2023Updated 3 years ago
- 🤖FFPA: Extends FA-2/3 via Split-D for large headdims, 1.5x~6×↑🎉 vs SDPA, up to 513~535 TFLOPS🎉 on NVIDIA H200.☆315Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.☆606May 13, 2026Updated 2 months ago
- A Quirky Assortment of CuTe Kernels☆1,063Updated this week
- MoE training for Me and You and maybe other people☆394Mar 15, 2026Updated 4 months ago
- Nsight Python is a Python kernel profiling interface based on NVIDIA Nsight Tools☆281Updated this week
- ☆253Nov 24, 2025Updated 7 months ago
- A PyTorch native platform for training generative AI models☆5,545Updated this week
- Scalable and Stable Parallelization of Nonlinear RNNS☆33Jun 28, 2026Updated 3 weeks ago