a whirlwind tour to deep learning and deep learning systems
☆81Aug 7, 2026Updated this week
Alternatives and similar repositories for asitpofborscht
Users that are interested in asitpofborscht are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Hand-Rolled GPU communications library☆96Nov 25, 2025Updated 8 months ago
- minimal compiler☆24Feb 19, 2026Updated 5 months ago
- NCCL communication API layer, and transport layer created from first principles.☆16Aug 20, 2025Updated 11 months ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆238Jun 29, 2026Updated last month
- The best ChatGPT that $100 can buy.☆58Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A blazing fast, MT-safe, lockfree and branchless circular byte buffer for SPSC in 50 loc☆13Sep 16, 2025Updated 10 months ago
- LLM RL envs done right, plus some training code☆17Feb 1, 2026Updated 6 months ago
- torch_remat fine-grained activation checkpointing API☆15Jul 28, 2026Updated last week
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated last month
- ☆19Dec 4, 2025Updated 8 months ago
- ☆22Jul 20, 2026Updated 2 weeks ago
- Code for the paper "Closing the Curious Case of Neural Text Degeneration"☆12Apr 9, 2025Updated last year
- MoE training for Me and You and maybe other people☆396Mar 15, 2026Updated 4 months ago
- Small scale distributed training of sequential deep learning models, built on Numpy and MPI.☆165Oct 19, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Course on Flash-attention in Triton☆100Feb 9, 2026Updated 6 months ago
- ☆21Mar 3, 2025Updated last year
- Batched square compact-Householder QR factorization.☆15Jul 2, 2026Updated last month
- ☆35Updated this week
- A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do☆1,286Apr 27, 2026Updated 3 months ago
- 🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models☆218May 27, 2026Updated 2 months ago
- ☆14Mar 29, 2026Updated 4 months ago
- An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.☆143Mar 21, 2026Updated 4 months ago
- 100M tokens. Infinite compute. Lowest val loss wins.☆522Jul 3, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is a beginner-friendly tutorial on MLIR from the perspective of a user of MLIR, not a compiler engineer. This tutorial will introduc…☆144Mar 5, 2026Updated 5 months ago
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆109Jul 29, 2026Updated last week
- Well documented examples of running distributed training jobs on Modal☆29Updated this week
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 10 months ago
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆142Sep 24, 2025Updated 10 months ago
- verl: Volcano Engine Reinforcement Learning for LLMs☆22Nov 6, 2025Updated 9 months ago
- Training framework with a goal to explore the frontier of sample efficiency of small language models☆101Jan 25, 2026Updated 6 months ago
- A curated list of reading material and lecture notes for all things geometry. Mostly focussed on differential and Riemannian geometry wit…☆23Jan 26, 2018Updated 8 years ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- GPU programming related news and material links☆2,260Jun 15, 2026Updated last month
- My solution for the LeHome challenge (1st place online, 2nd place in real world round)☆87Jun 26, 2026Updated last month
- ☆68Apr 8, 2026Updated 4 months ago
- Code for "Approaching Deep Learning through the Spectral Dynamics of Weights"☆13Oct 30, 2024Updated last year
- ☆10Dec 17, 2019Updated 6 years ago
- Educational WIP☆73Feb 16, 2026Updated 5 months ago
- Agentic RL Training at Scale☆1,855Updated this week