Implement FlashAttention v2 with minimal code to learn.
☆19Jun 12, 2024Updated 2 years ago
Alternatives and similar repositories for flash-attention-v2-minimal
Users that are interested in flash-attention-v2-minimal are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash Attention in raw Cuda C beating PyTorch☆39May 14, 2024Updated 2 years ago
- NES emulator written in pure FreeBASIC with love by Blyss Sarania and Gavin Schulte(Nobbs66).☆21Oct 29, 2025Updated 9 months ago
- ☆10Mar 14, 2018Updated 8 years ago
- 🤖 Telegram chatbot frontend for Searx.☆16Nov 25, 2018Updated 7 years ago
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Nov 25, 2019Updated 6 years ago
- A rust version of the Caffe library.☆19Jun 16, 2021Updated 5 years ago
- ☆17Apr 9, 2025Updated last year
- ☆32Feb 20, 2024Updated 2 years ago
- ☆15Apr 28, 2023Updated 3 years ago
- VGG16 architecture with BatchNorm☆14Apr 4, 2017Updated 9 years ago
- ☆14Jul 7, 2026Updated last month
- NetHCF: Enabling Line-rate and Adaptive Spoofed IP Traffic Filtering☆13Mar 17, 2022Updated 4 years ago
- ☆14Dec 3, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A CUDA kernel for NHWC GroupNorm for PyTorch☆26Nov 15, 2024Updated last year
- 关于深度学习算法、框架、编译器、加速器的一些理解☆16Jul 2, 2022Updated 4 years ago
- [NeurIPS 2024] The official implementation of ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification☆33Mar 30, 2025Updated last year
- A graph coloring register allocator for LLVM.☆11Jan 23, 2017Updated 9 years ago
- a PureBasic FrameWork☆13Mar 3, 2026Updated 5 months ago
- ☆17Apr 30, 2025Updated last year
- Acclaim: Adaptive Memory Reclaim to Improve User Experience in Android Systems [ATC '20]☆16Aug 1, 2020Updated 6 years ago
- seeta face detection for Android☆11Sep 23, 2017Updated 8 years ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Immix GC for LLVM based languages☆17Apr 2, 2025Updated last year
- 使用 cutlass 实现 flash-attention 精简版,具有教学意义☆59Aug 12, 2024Updated 2 years ago
- Visualize the Expert Parallelism Load Balancer☆19Mar 15, 2025Updated last year
- flash attention tutorial written in python, triton, cuda, cutlass☆532Jan 20, 2026Updated 6 months ago
- MLIR backend for optimising graph algorithms☆17Mar 30, 2024Updated 2 years ago
- Optimizing Tensor Computation Graphs with Equality Saturation and Monte Carlo Tree Search☆15Aug 9, 2024Updated 2 years ago
- Work in progress LLM framework.☆16Oct 31, 2024Updated last year
- A super simple web interface to perform blind tests on LLM outputs.☆30Mar 9, 2024Updated 2 years ago
- Tender: Accelerating Large Language Models via Tensor Decompostion and Runtime Requantization (ISCA'24)☆34Jul 4, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA en…☆54Aug 25, 2024Updated last year
- Simple compiler that translates simpified C into ARMv7 assembly☆19Apr 12, 2024Updated 2 years ago
- 自动查 [SDU 青岛校区] 宿舍电量,低于阈值则邮件提醒(Github Actions,Serverless)☆16May 26, 2021Updated 5 years ago
- XREAL Visual-Inertial Odometry for 6DOF VR☆16Jun 29, 2025Updated last year
- laboratory assignments of cs143-Compilers☆18Jun 8, 2021Updated 5 years ago
- Get down and dirty with FlashAttention2.0 in pytorch, plug in and play no complex CUDA kernels☆115Jul 31, 2023Updated 3 years ago
- ☆28Oct 25, 2021Updated 4 years ago