Small scale distributed training of sequential deep learning models, built on Numpy and MPI.
☆166Oct 19, 2023Updated 2 years ago
Alternatives and similar repositories for ShallowSpeed
Users that are interested in ShallowSpeed are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆93Jul 5, 2024Updated 2 years ago
- Fast CUDA matrix multiplication from scratch☆1,295Sep 2, 2025Updated 11 months ago
- Step-by-step optimization of CUDA SGEMM☆497Mar 30, 2022Updated 4 years ago
- Minimal but scalable implementation of large language models in JAX☆34Nov 28, 2025Updated 9 months ago
- a whirlwind tour to deep learning and deep learning systems☆82Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆591Jul 11, 2024Updated 2 years ago
- Fastest kernels written from scratch☆615Aug 15, 2026Updated 2 weeks ago
- GPU programming related news and material links☆2,311Jun 15, 2026Updated 2 months ago
- What would you do with 1000 H100s...☆1,191Jan 10, 2024Updated 2 years ago
- 🔀 yet another mixture of experts☆23Jun 5, 2026Updated 2 months ago
- a tiny vectorstore implementation built with numpy.☆64Apr 26, 2024Updated 2 years ago
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,292Aug 26, 2025Updated last year
- ☆62Feb 24, 2026Updated 6 months ago
- ☆30Dec 2, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆46May 24, 2025Updated last year
- writing really fast kernels☆19Jul 15, 2026Updated last month
- Algorithms for approximate attention in LLMs☆22Apr 14, 2025Updated last year
- ☆172Dec 27, 2024Updated last year
- coloring terminal text with intensities (used for plotting probability, entropy with tokens)☆12Oct 11, 2024Updated last year
- ring-attention experiments☆174Oct 17, 2024Updated last year
- A repository to unravel the language of GPUs, making their kernel conversations easy to understand☆213Jun 1, 2025Updated last year
- UNet diffusion model in pure CUDA☆663Jun 28, 2024Updated 2 years ago
- Tile primitives for speedy kernels☆3,659Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Large scale 4D parallelism pre-training for 🤗 transformers in Mixture of Experts *(still work in progress)*☆87Dec 14, 2023Updated 2 years ago
- Puzzles for learning Triton☆2,577Apr 1, 2026Updated 4 months ago
- ☆51Jan 18, 2024Updated 2 years ago
- run paligemma in real time☆132May 18, 2024Updated 2 years ago
- Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.☆329Updated this week
- A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.☆604Aug 14, 2026Updated 2 weeks ago
- A Quirky Assortment of CuTe Kernels☆1,133Aug 21, 2026Updated last week
- MoE training for Me and You and maybe other people☆396Mar 15, 2026Updated 5 months ago
- Nsight Python is a Python kernel profiling interface based on NVIDIA Nsight Tools☆289Aug 17, 2026Updated 2 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆258Nov 24, 2025Updated 9 months ago
- A PyTorch native platform for training generative AI models☆5,680Updated this week
- Scalable and Stable Parallelization of Nonlinear RNNS☆33Jun 28, 2026Updated 2 months ago
- LLM training in simple, raw C/CUDA☆114May 1, 2024Updated 2 years ago
- ☆355Updated this week
- ☆190Jun 16, 2024Updated 2 years ago
- 삼각형의 실전! Triton☆16Feb 15, 2024Updated 2 years ago