forked from vllm-project/flash-attention
☆64May 9, 2026Updated 4 months ago
Alternatives and similar repositories for flash-attention-v100
Users that are interested in flash-attention-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 2 months ago
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆206Jun 30, 2026Updated 2 months ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 5 months ago
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,010Updated this week
- vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support☆23Jun 20, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆452Updated this week
- Local LLM Inference Speed Test Tool☆215Sep 2, 2026Updated last week
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- golang live - rtmp - httpflv - hls☆11Jul 10, 2026Updated 2 months ago
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆136Updated this week
- Run DeepSeek-V4.1-Flash / DeepSeek-V4-Flash and GLM-5.3-Flash on SM89 (Ada / RTX 4090) and SM120 (RTX PRO 6000) with vLLM☆157Updated this week
- File source using external command for ddu.vim☆21Jul 20, 2026Updated last month
- 我的nonebot机器人☆15Sep 7, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Rust-based, SenseVoiceSmall☆40Updated this week
- CI scripts designed to build a Pascal-compatible version of vLLM.☆13Aug 10, 2024Updated 2 years ago
- A proof of concept tool for using local LLMs to transform messy text documents into structured JSON☆26Aug 22, 2024Updated 2 years ago
- ☆17Oct 15, 2023Updated 2 years ago
- 引流外链【开心版引流宝】致力于为个人、团队提供基于微信私域流量的推广、引流的效率工具。可减轻人力,有效降低资源损失、流量流失的几率。引流宝完全开源,免费,可商用、可任意二次开发。引流宝可以辅助你更好地开展营销活动推广!降低运营成本,提高工作效率,获取更多资源。☆13Feb 12, 2024Updated 2 years ago
- Revision of official yolov7-pose to support custom dataset for keypoint detection☆11Nov 12, 2023Updated 2 years ago
- ☆16Oct 11, 2025Updated 11 months ago
- A simple, cross-platform CLI tool for quickly switching between Claude Code configuration profiles by managing different settings.json ve…☆17Jan 17, 2026Updated 7 months ago
- 原 XiunoBBS 自带的插件☆21Apr 10, 2022Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- stable diffusion running on Google TPU.☆13Sep 7, 2022Updated 4 years ago
- ☆26Jul 30, 2026Updated last month
- A Cursor Agent Skill that turns math problems (esp. geometry) into faithful + interactive GeoGebra .ggb files by driving the real GeoGebr…☆63Jul 2, 2026Updated 2 months ago
- DeepSiteForOpenAI 是一个灵活的开发工具,它将 DeepSite 的强大功能与 OpenAI 接口无缝集成,支持自定义接口,使用openai 风格的接口为开发者提供了一个高效、智能的编程环境。这个工具允许用户通过自然语言描述来生成代码,实现"氛围编程"(Vi…☆15Apr 9, 2025Updated last year
- triton3.2.0添加mi25/mi50/mi60支持☆14Apr 26, 2025Updated last year
- Dogfooding vLLM backport for older GPUs.☆213Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,221Updated this week
- ChatAdmin是ChatFlow和ChatStudio的后端API服务☆18Oct 18, 2024Updated last year
- The official pytorch implementation of “Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization”.☆19May 22, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- WPS Office 全家桶(文字/表格/演示/PDF)AI 助手:挂 OpenAI 兼容 / Anthropic / Codex 多家模型,AI 通过工具调用直接读写文档;支持生图、MCP、中英双语,覆盖 Windows / macOS / Linux(含国产系统)。☆52Aug 11, 2026Updated last month
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆435Feb 20, 2026Updated 6 months ago
- Use yolov5 to realize the road occupation operation and vehicle parking violation detection in urban streets, and can independently delin…☆13Jan 2, 2023Updated 3 years ago
- Workflow-to-APP、ScreenShare&FloatingVideo、GPT & 3D、SpeechRecognition&TTS☆42May 11, 2026Updated 4 months ago
- A ComfyUI custom node implementation of TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows.☆45Mar 6, 2026Updated 6 months ago
- Filezilla for Windows on ARM☆22Mar 21, 2022Updated 4 years ago
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year