small auto-grad engine inspired from Karpathy's micrograd and PyTorch
☆276Nov 21, 2024Updated last year
Alternatives and similar repositories for smolgrad
Users that are interested in smolgrad are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- a tiny multidimensional array implementation in C similar to numpy, but only one file.☆224Aug 2, 2024Updated 2 years ago
- Learnings and programs related to CUDA☆438Jun 29, 2025Updated last year
- a tiny vectorstore implementation built with numpy.☆64Apr 26, 2024Updated 2 years ago
- learningggggggg 🐳☆639Apr 2, 2025Updated last year
- A repository consisting of paper/architecture replications of classic/SOTA AI/ML papers in pytorch☆428Nov 11, 2025Updated 10 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- High Quality Resources on GPU Programming/Architecture☆593Jul 26, 2024Updated 2 years ago
- A zero-dependency ML framework in C with a modern Python API for full control over execution and memory.☆707Updated this week
- This repo has all the basic things you'll need in-order to understand complete vision transformer architecture and its various implementa…☆228Jan 2, 2025Updated last year
- Zig tensor library☆24Jul 11, 2025Updated last year
- From the Tensor to Stable Diffusion, a rough outline for a 10 week course.☆1,082Apr 5, 2026Updated 5 months ago
- Andrej Kapathy's micrograd implemented in c☆30Aug 7, 2024Updated 2 years ago
- MLX port for xjdr's entropix sampler (mimics jax implementation)☆62Nov 4, 2024Updated last year
- idk☆24Jun 7, 2025Updated last year
- Simple Transformer in Jax☆143Jun 22, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A deep-dive on the entire history of deep-learning☆1,580Jul 16, 2024Updated 2 years ago
- parallelized hyperdimensional tictactoe☆128Aug 25, 2024Updated 2 years ago
- An ML Systems Onboarding list☆1,134Feb 19, 2026Updated 7 months ago
- An AI Powered Planner and Time Management web app☆43May 29, 2024Updated 2 years ago
- my little linear algebra library☆43Jul 7, 2024Updated 2 years ago
- A curated list of resources for learning and exploring Triton, OpenAI's programming language for writing efficient GPU code.☆498Mar 10, 2025Updated last year
- Flash Attention from scratch, tiled CUDA forward kernel, online softmax with running max and correction factor, recomputation trick in ba…☆19Mar 6, 2026Updated 6 months ago
- links for research internships☆156Jan 28, 2025Updated last year
- Hand-Rolled GPU communications library☆94Nov 25, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- machine learning from absolute scratch in c. gradients, linear algebra ops & everything else without using any third party library!☆26Aug 3, 2024Updated 2 years ago
- A MNIST neural network written from scratch in Odin, visualised with Raylib☆181Sep 25, 2024Updated 2 years ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- Tensor library with autograd using only Rust's standard library☆73Jul 1, 2024Updated 2 years ago
- simply neural networks in every language☆42Nov 24, 2025Updated 10 months ago
- UNet diffusion model in pure CUDA☆666Jun 28, 2024Updated 2 years ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆63Feb 25, 2026Updated 7 months ago
- Basically a repo containing architectures/algorithms/papers from scratch in pytorch☆30Feb 11, 2026Updated 7 months ago
- Multi-Threaded FP32 Matrix Multiplication on x86 CPUs☆377Apr 21, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Whalegrad 🐳 is a lightweight deep learning library written in C.☆10Jan 5, 2025Updated last year
- This is a small autograd engine, made purely from numpy and python.☆27Sep 17, 2024Updated 2 years ago
- NanoGPT (124M) in 90 seconds☆5,831Sep 18, 2026Updated last week
- Learning about CUDA by writing PTX code.☆162Feb 27, 2024Updated 2 years ago
- A 120-day CUDA learning plan covering daily concepts, exercises, pitfalls, and references (including “Programming Massively Parallel Proc…☆962Mar 29, 2025Updated last year
- FlexAttention based, minimal vllm-style inference engine for fast Gemma 2 inference.☆360Nov 2, 2025Updated 10 months ago
- python driver and runtime for tenstorrent blackhole cards☆20Sep 16, 2026Updated last week