Fast and memory-efficient exact attention
☆20Jul 22, 2024Updated 2 years ago
Alternatives and similar repositories for flash-attention
Users that are interested in flash-attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simple implementation of a deep linear Pytorch module☆21Oct 16, 2020Updated 5 years ago
- A simple cycle-accurate DaDianNao simulator☆13Mar 27, 2019Updated 7 years ago
- Implementation of N-Grammer, augmenting Transformers with latent n-grams, in Pytorch☆82Dec 4, 2022Updated 3 years ago
- Graph neural network message passing reframed as a Transformer with local attention☆70Dec 24, 2022Updated 3 years ago
- Implementation of Cross Transformer for spatially-aware few-shot transfer, in Pytorch☆54Mar 30, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆13Apr 16, 2022Updated 4 years ago
- Maximal Update Parametrization (μP) with Flax & Optax.☆16Dec 27, 2023Updated 2 years ago
- Hardware Division Units☆10Jul 17, 2014Updated 12 years ago
- ☆10May 24, 2020Updated 6 years ago
- RADIX-4 SRT division☆13Oct 31, 2019Updated 6 years ago
- Radam+lookahead implemented by tensorflow☆11Oct 14, 2019Updated 6 years ago
- ☆12Jun 12, 2017Updated 9 years ago
- Code for H. Narasimhan, "Learning with Complex Loss Functions and Constraints", AISTATS 2018☆11Mar 21, 2018Updated 8 years ago
- Learning Transferable Features with Deep Adaptation Networks☆12Jul 18, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 🎓 The chrome extension to make learning from YouTube faster & easier.☆11Jan 9, 2022Updated 4 years ago
- Diagram-first LLM architecture, memory, hardware-fit, throughput, and verified cloud-cost explorer.☆22Updated this week
- Pytorch implementation of "Very Deep Graph Neural Networks via Noise Regularisation"☆10Aug 22, 2021Updated 5 years ago
- ☆15Jun 26, 2023Updated 3 years ago
- ☆13Jun 4, 2024Updated 2 years ago
- tensorflow implementation for scoring blur image sharpness☆12Nov 29, 2017Updated 8 years ago
- A pipeline for the automatic construction of geometry problems along with step-by-step solutions.☆18Aug 27, 2025Updated last year
- An Agile RISC-V SoC Design Framework with in-order cores, out-of-order cores, accelerators, and more☆13May 29, 2026Updated 3 months ago
- A simple cross attention that updates both the source and target in one step☆199Jul 29, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Implementation of "Denoise Pretraining on Non-equilibrium Molecular Conformations for Accurate and Transferable Neural Potentials" in PyT…☆14Jul 26, 2023Updated 3 years ago
- NoC based MPSoC☆11Jul 17, 2014Updated 12 years ago
- (Verilog) A simple convolution layer implementation with systolic array structure☆14May 9, 2022Updated 4 years ago
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago
- A PyTorch implementation of [VCT](https://github.com/google-research/google-research/tree/master/vct)☆10Nov 25, 2022Updated 3 years ago
- Implementation of Recurrent Independent Mechanisms in Pytorch☆27Apr 6, 2026Updated 4 months ago
- Basic floating-point components for RISC-V processors☆12Aug 13, 2017Updated 9 years ago
- Showing full TensorBoard support in Tensorflow for a CNN using MNIST data.☆13Oct 19, 2019Updated 6 years ago
- Associative scan package for DRYing some code between repos☆19Jul 30, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A implement of run-length encoding for Pytorch tensor using CUDA☆13Apr 7, 2021Updated 5 years ago
- Implementation of numerous Vision Transformers in Google's JAX and Flax.☆22Aug 30, 2022Updated 4 years ago
- 本科编译原理大作业:Verilog to Python Testbench Module:生成 FIRRTL 中间表示的 Verilog 文法子集的前端与基于 Arcilator 生成 Python 仿真模块的后端☆16Jan 8, 2025Updated last year
- unsigned Radix-2 SRT division,基2除法☆17May 12, 2015Updated 11 years ago
- Exploring Motion Ambiguity and Alignment for High-Quality Video Frame Interpolation (CVPR2023)☆14Jul 21, 2023Updated 3 years ago
- ☆15Jan 27, 2025Updated last year
- Complex-Edit: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark☆29Apr 22, 2025Updated last year