a whirlwind tour to deep learning and deep learning systems
☆82Aug 27, 2026Updated this week
Alternatives and similar repositories for asitpofborscht
Users that are interested in asitpofborscht are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Hand-Rolled GPU communications library☆96Nov 25, 2025Updated 9 months ago
- minimal compiler☆24Feb 19, 2026Updated 6 months ago
- ☆22Apr 22, 2024Updated 2 years ago
- NCCL communication API layer, and transport layer created from first principles.☆16Aug 20, 2025Updated last year
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆240Jun 29, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- LLM RL envs done right, plus some training code☆17Feb 1, 2026Updated 6 months ago
- torch_remat fine-grained activation checkpointing API☆16Aug 19, 2026Updated last week
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated last month
- ☆19Dec 4, 2025Updated 8 months ago
- ☆22Jul 20, 2026Updated last month
- Code for the paper "Closing the Curious Case of Neural Text Degeneration"☆12Apr 9, 2025Updated last year
- MoE training for Me and You and maybe other people☆396Mar 15, 2026Updated 5 months ago
- Small scale distributed training of sequential deep learning models, built on Numpy and MPI.☆166Oct 19, 2023Updated 2 years ago
- Course on Flash-attention in Triton☆102Feb 9, 2026Updated 6 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- minimal DL library in C: 24 NAIVE cuda/cpu ops, autodiff engine, python API (ops bindings/layers/models), tensor abstraction, strides, co…☆62Dec 17, 2025Updated 8 months ago
- ☆21Mar 3, 2025Updated last year
- Batched square compact-Householder QR factorization.☆15Jul 2, 2026Updated last month
- ☆43Aug 16, 2026Updated 2 weeks ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆2,183Aug 23, 2026Updated last week
- 🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models☆223May 27, 2026Updated 3 months ago
- ☆14Mar 29, 2026Updated 5 months ago
- An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.☆144Mar 21, 2026Updated 5 months ago
- 100M tokens. Infinite compute. Lowest val loss wins.☆528Jul 3, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆354Updated this week
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆112Jul 29, 2026Updated last month
- Code for verifying deep neural feature ansatz☆23May 3, 2023Updated 3 years ago
- In fluid dynamics, an eddy is the swirling of a fluid and the reverse current created when the fluid is in a turbulent flow regime.☆19Jul 23, 2026Updated last month
- Well documented examples of running distributed training jobs on Modal☆33Aug 11, 2026Updated 2 weeks ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 11 months ago
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆142Aug 15, 2026Updated 2 weeks ago
- verl: Volcano Engine Reinforcement Learning for LLMs☆23Nov 6, 2025Updated 9 months ago
- Training framework with a goal to explore the frontier of sample efficiency of small language models☆101Jan 25, 2026Updated 7 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A curated list of reading material and lecture notes for all things geometry. Mostly focussed on differential and Riemannian geometry wit…☆23Jan 26, 2018Updated 8 years ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 6 months ago
- GPU programming related news and material links☆2,310Jun 15, 2026Updated 2 months ago
- Summer of Math Exposition☆40Updated this week
- My solution for the LeHome challenge (1st place online, 2nd place in real world round)☆101Aug 16, 2026Updated 2 weeks ago
- ☆109Aug 13, 2026Updated 2 weeks ago
- Code for "Approaching Deep Learning through the Spectral Dynamics of Weights"☆13Oct 30, 2024Updated last year