☆78Feb 19, 2024Updated 2 years ago
Alternatives and similar repositories for flash-attention-v100
Users that are interested in flash-attention-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash Attention in raw Cuda C beating PyTorch☆39May 14, 2024Updated 2 years ago
- 使用 cutlass 实现 flash-attention 精简版,具有教学意义☆59Aug 12, 2024Updated 2 years ago
- LV-BERT: Exploiting Layer Variety for BERT (Findings of ACL 2021)☆19May 10, 2023Updated 3 years ago
- Fast and low-memory attention layer written in CUDA☆20Jul 14, 2023Updated 3 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆46Feb 27, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆14Dec 22, 2024Updated last year
- A vector field rendering library☆17Jul 31, 2019Updated 7 years ago
- LLM inference in C/C++☆30Updated this week
- High-performance GEMM implementation optimized for NVIDIA H100 GPUs, leveraging Hopper architecture's TMA, WGMMA, and Thread Block Cluste…☆11Aug 25, 2026Updated 3 weeks ago
- ☆12Mar 21, 2024Updated 2 years ago
- 2014: Variational Monte Carlo for the harmonic oscillator, helium, hydrogen and H2 - IPython notebook and FORTRAN90☆13Jun 23, 2016Updated 10 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆15Aug 31, 2023Updated 3 years ago
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- Official code for "Shapley Values-enabled Progressive Pseudo Bag Augmentation for Whole-Slide Image Classification", IEEE Transaction on …☆14Jul 9, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- 本项目分享了本人在四川大学计算机学院计算机科学与技术专业的各类课程的资料、学习建议以及作业。欢迎使用,也希望其他校友能为此库提供缺失资料,如果喜欢就Star吧。☆10May 18, 2021Updated 5 years ago
- writing really fast kernels☆20Sep 12, 2026Updated last week
- Code for Unsupervised multi-granular Chinese word segmentation and term discovery via graph partition [JBI]☆16Jan 28, 2022Updated 4 years ago
- Python爬虫获取QQ群精华消息☆17Aug 7, 2024Updated 2 years ago
- forked from vllm-project/flash-attention☆65May 9, 2026Updated 4 months ago
- ☆79Dec 15, 2023Updated 2 years ago
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- ☆12Jul 20, 2022Updated 4 years ago
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,102Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated 2 years ago
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- Pointer analysis prototype (currently including anderson, steensgard).☆17Dec 20, 2021Updated 4 years ago
- ☆14Sep 4, 2024Updated 2 years ago
- Jetbrains plugin for codetime☆16Apr 3, 2023Updated 3 years ago
- Benchmarks for High-Level Synthesis☆11Mar 17, 2023Updated 3 years ago
- UESTC-《Parallel and Distributed Computing》Course Experiment(电子科技大学 《分布式并行计算》课程实验)-Nvidia CUDA Course on https://courses.nvidia.com/course…☆10Jun 18, 2019Updated 7 years ago
- Flash Attention in ~100 lines of CUDA (forward pass only)☆1,186Dec 30, 2024Updated last year
- 四川大学硕博学位论文LaTeX模板v1.0 beta版本 power by scuthesis of LegendaryLeo☆13Apr 4, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆11May 5, 2021Updated 5 years ago
- ☆17Jul 18, 2026Updated 2 months ago
- An out-of-the-box inference acceleration engine for Diffusion and DiT models☆59Mar 21, 2025Updated last year
- 基于中文 GPT2 预训练模型的语句困惑度计算☆15Apr 20, 2023Updated 3 years ago
- Bloom filter alternative (C++)☆18Nov 8, 2018Updated 7 years ago
- Groundhog - Serial ATA Host Bus Adapter☆24Jun 10, 2018Updated 8 years ago
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago