The Little Book of Reinforcement Learning
☆1,610Jul 14, 2026Updated 2 months ago
Alternatives and similar repositories for little-book-rl
Users that are interested in little-book-rl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated list of best cuda programming books☆984May 19, 2026Updated 4 months ago
- Textbook on reinforcement learning from human feedback☆2,404Sep 11, 2026Updated last week
- RL-training an AI agent to RL-train AI agents.☆239Jul 14, 2026Updated 2 months ago
- A concise, beginner-friendly introduction to the core ideas of linear algebra.☆2,001Mar 16, 2026Updated 6 months ago
- Speedrunning LoRA fine-tuning: frozen task, frozen hardware, public wall-clock leaderboard. modded-nanogpt for fine-tuning.☆148Sep 12, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."☆17,864Updated this week
- A ~9M parameter LLM that talks like a small fish.☆3,771Apr 15, 2026Updated 5 months ago
- GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.☆121Jun 18, 2026Updated 3 months ago
- A library for incremental computations☆1,514Jul 10, 2026Updated 2 months ago
- Run models too big for your Mac's memory☆668Aug 27, 2026Updated 3 weeks ago
- Machine Learning Engineering Open Book☆19,031Sep 12, 2026Updated last week
- The best ChatGPT that $100 can buy.☆58,214Sep 7, 2026Updated 2 weeks ago
- A theoretical and practical deep dive into Reinforcement Learning with Human Feedback and it’s applications in Large Language Models from…☆117Nov 7, 2025Updated 10 months ago
- A minimal hackable implementation of policy gradient methods (GRPO, PPO, REINFORCE)☆18Feb 20, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference…☆687Apr 16, 2026Updated 5 months ago
- Voxtral ASR & TTS running natively and in the browser. A Rust implementation of Mistral's Voxtral mini realtime ASR / TTS using the Burn …☆822Apr 2, 2026Updated 5 months ago
- ☆3,413May 5, 2026Updated 4 months ago
- Implement a ChatGPT-like LLM in PyTorch from scratch, step by step☆105,393Updated this week
- Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.☆629Feb 24, 2025Updated last year
- tiny torch, but close to metal☆134Jun 25, 2026Updated 2 months ago
- Machine Learning Systems: Foundations, Scaling, Agentic AI, and Physical AI (Vols I–IV) • Harvard CS249r | https://mlsysbook.ai☆28,411Updated this week
- A reimplementation of Stable Diffusion 3.5 in pure PyTorch☆708Jun 14, 2025Updated last year
- Deep Reinforcement Learning: Zero to Hero!☆2,296May 26, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 5 months ago
- Sutskever 30 implementations inspired by https://papercode.vercel.app/ | For Agents, use https://github.com/pageman/Sutskever-Agent | Pol…☆4,564Mar 15, 2026Updated 6 months ago
- Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM☆1,128Sep 15, 2026Updated last week
- ☆64Jul 17, 2026Updated 2 months ago
- R.L. methods and techniques.☆197Sep 9, 2026Updated 2 weeks ago
- A concise book on quantum mechanics, intended for anyone who knows linear algebra. Planning to add a chapter on entanglement, and appendi…☆156Jul 7, 2026Updated 2 months ago
- Browser-based Quake III Arena map editor with WebGL2 and client-side BSP compilation.☆21Aug 5, 2026Updated last month
- A hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exe…☆193Jul 17, 2026Updated 2 months ago
- A generative pretrained transformer implementation☆94Aug 30, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LLM101n: Let's build a Storyteller☆37,490Aug 1, 2024Updated 2 years ago
- Everything I know about running LLMs locally☆1,842Jul 10, 2026Updated 2 months ago
- Distributed DuckDB - dual execution and differential storage☆571May 6, 2026Updated 4 months ago
- 🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.☆4,429Sep 3, 2026Updated 2 weeks ago
- Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smar…☆12,158Updated this week
- ☆7,232Aug 14, 2026Updated last month
- Implement a reasoning LLM in PyTorch from scratch, step by step☆5,272Updated this week