forked from vllm-project/flash-attention
☆66May 9, 2026Updated 4 months ago
Alternatives and similar repositories for flash-attention-v100
Users that are interested in flash-attention-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 3 months ago
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆212Sep 22, 2026Updated last week
- 1CatV2 with TileLANG written FA-v100 and many goodies☆19Jul 15, 2026Updated 2 months ago
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,220Updated this week
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆30Feb 9, 2026Updated 7 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 10 months ago
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆462Sep 22, 2026Updated last week
- Local LLM Inference Speed Test Tool☆231Sep 2, 2026Updated last month
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆81Sep 19, 2026Updated 2 weeks ago
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- This repository is created to provide a straight forward reference for running Custom Yolov5 models on the NPU for Orangepi 5 boards - RK…☆13Feb 19, 2025Updated last year
- Run DeepSeek-V4.1-Flash / DeepSeek-V4-Flash and GLM-5.3-Flash on SM89 (Ada / RTX 4090) and SM120 (RTX PRO 6000) with vLLM☆177Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆142Updated this week
- Modified for AMD MI50 GPUs | A high-throughput and memory-efficient inference and serving engine for LLMs☆15Mar 2, 2026Updated 7 months ago
- TensorDB: In-Database Tensor Manipulation with Tensor-Relational Query Plans☆21Jul 25, 2014Updated 12 years ago
- Qt sample app that demonstrates the QSyntaxHighlighter class☆12Feb 19, 2020Updated 6 years ago
- CNN训练与测试人脸戴眼镜与否的图片分类(TensorFlow)☆30Dec 8, 2017Updated 8 years ago
- ☆79Feb 19, 2024Updated 2 years ago
- ☆23Dec 15, 2016Updated 9 years ago
- ☆17Oct 15, 2023Updated 2 years ago
- The complete NUMA-optimized branch of the ktransformers project☆26Nov 3, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- NoneBot 黑白名单☆15May 7, 2023Updated 3 years ago
- Revision of official yolov7-pose to support custom dataset for keypoint detection☆11Nov 12, 2023Updated 2 years ago
- 模拟 firefox 请求☆24Jun 30, 2026Updated 3 months ago
- A simple, cross-platform CLI tool for quickly switching between Claude Code configuration profiles by managing different settings.json ve…☆18Jan 17, 2026Updated 8 months ago
- ImageJ/Fiji plugin for consistent elastic registration of 2D images☆26Jun 14, 2022Updated 4 years ago
- 适用于hoshino的冰祈插件集合☆21Jun 15, 2023Updated 3 years ago
- 使用Live2D与GPT-sovits的AI全自动直播一站式解决方案。☆19Jun 10, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,274Updated this week
- bge-large-zh api service☆24Jan 5, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A distributed actor-based framework built on Microsoft Orleans for building scalable event-sourced applications.☆49Dec 17, 2025Updated 9 months ago
- PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM☆36Apr 13, 2026Updated 5 months ago
- The official pytorch implementation of “Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization”.☆19May 22, 2025Updated last year
- 500彩票网球赛数据爬取:通过实时爬取足球赛事信息,赛事胜平负及欧赔率及凯利指数,计算变异系数,图表展示。☆50Feb 15, 2023Updated 3 years ago
- WPS Office 全家桶(文字/表格/演示/PDF)AI 助手:挂 OpenAI 兼容 / Anthropic / Codex 多家模型,AI 通过工具调用直接读写文档;支持生图、MCP、中英双语,覆盖 Windows / macOS / Linux(含国产系统)。☆65Aug 11, 2026Updated last month
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 7 months ago
- Use yolov5 to realize the road occupation operation and vehicle parking violation detection in urban streets, and can independently delin…☆13Jan 2, 2023Updated 3 years ago