Flash Attention from scratch, tiled CUDA forward kernel, online softmax with running max and correction factor, recomputation trick in backward, O(N) memory, full forward and backward verified against PyTorch autograd to 1e-6.
☆19Mar 6, 2026Updated 7 months ago
Alternatives and similar repositories for FlashAttention-CuPy
Users that are interested in FlashAttention-CuPy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆12Oct 28, 2023Updated 2 years ago
- a minimal paged attention implementation☆20Jan 30, 2026Updated 8 months ago
- Code for the paper "FedFisher: Leveraging Fisher Information for One-Shot Federated Learning" by Divyansh Jhunjhunwala, Shiqiang Wang, an…☆15Feb 10, 2024Updated 2 years ago
- The best ChatGPT that $100 can buy.☆59Updated this week
- Integrates Imbue's Cost Aware pareto-Region Bayesian Search (CARBS) with Weights and Biases (WanDB)☆12Mar 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆44May 17, 2026Updated 4 months ago
- ☆27Sep 9, 2024Updated 2 years ago
- Interactive version of the CuTe layout paper☆57Apr 14, 2026Updated 5 months ago
- [ICLR23] First deep learning-based surrogate model that jointly learns the evolution model and optimizes computational cost via remeshing☆42Jan 30, 2024Updated 2 years ago
- This is a simple automated license plate detector developed in C++ via OpenCV.☆11Sep 26, 2020Updated 6 years ago
- ☆14Feb 9, 2026Updated 7 months ago
- Mini RL Lab☆16Jun 17, 2024Updated 2 years ago
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Oct 1, 2026Updated last week
- ☆292Sep 30, 2026Updated last week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Improving AI Systems with Self-Defense Mechanisms☆25Feb 28, 2025Updated last year
- clear single-file JAX implementations of common RL algorithms☆15Sep 5, 2021Updated 5 years ago
- Adaptation of titans-pytorch to llama models on HF☆26Mar 6, 2025Updated last year
- ☆13May 25, 2023Updated 3 years ago
- AI-queryable MCP server for unified real-time air, sea, and space tracking — 25+ tools, agent skills, guardrails, and a 3D globe☆18Apr 15, 2026Updated 5 months ago
- ☆18Aug 14, 2023Updated 3 years ago
- A simple but instructive implementation of DP, TP, FSDP, FSDP+TP using pytorch distributed primitives☆24Apr 12, 2026Updated 5 months ago
- A NodeJS application to upload, watch and stream live videos.☆12Jan 24, 2023Updated 3 years ago
- Rust FTL + WebRTC live streaming software.☆13Mar 12, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- CartPole-v0 via PPO with GAE, PyTorch☆22Feb 10, 2019Updated 7 years ago
- Adapting LLaMA Decoder to Vision Transformer☆31May 20, 2024Updated 2 years ago
- Website for the ICML 2021 tutorial on Random Matrix Theory and Machine Learning☆16Dec 8, 2021Updated 4 years ago
- Brax + Pufferlib + CARBS for gpu-accelerated robotics RL☆12Jun 12, 2025Updated last year
- Meta in-context learning for protein fitness prediction☆21Feb 7, 2025Updated last year
- Microbenchmarking hyperparameter tuning for JAX functions.☆24Sep 24, 2026Updated last week
- Algorithms for Gradient TD updates☆19Feb 21, 2026Updated 7 months ago
- Pre-flight checks for PyTorch pipelines. Catch silent failures before they waste your GPU.☆24Mar 15, 2026Updated 6 months ago
- [ 👾 ] ➡️ 💾 ➡️ { 🎮🕹️ } Extra Stable-Baselines3 buffer classes. Reducing RL memory usage drastically with minimal overhead.☆24Jun 9, 2026Updated 3 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- WhisperMesh is an advanced chatbot that integrates voice and text interactions, delivering personalized responses through LLM models and …☆17Apr 23, 2025Updated last year
- The GPT-4 function calls used in everchanging quest for the HF game jam☆10Jul 9, 2023Updated 3 years ago
- A Mac App to find unused source files in XCode project.☆13Sep 24, 2018Updated 8 years ago
- ARIMA model from scratch using numpy and pandas.☆25Jul 14, 2021Updated 5 years ago
- Evaluation Code repository for the paper "ModuLoRA: Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers". (2023…☆13Dec 5, 2023Updated 2 years ago
- RL models to play Sokoban. The fastest recipe wins.☆33Jul 25, 2026Updated 2 months ago
- Convert gltf or glb files to fbx using a blender 2.8 docker image.☆16Feb 13, 2019Updated 7 years ago