forked from vllm-project/flash-attention
☆62May 9, 2026Updated 3 months ago
Alternatives and similar repositories for flash-attention-v100
Users that are interested in flash-attention-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆194Jun 30, 2026Updated last month
- 1CatV2 with TileLANG written FA-v100 and many goodies☆18Jul 15, 2026Updated last month
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆447Aug 19, 2026Updated last week
- Local LLM Inference Speed Test Tool☆177Jul 30, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆71May 4, 2025Updated last year
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Jul 14, 2026Updated last month
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 8 months ago
- galge.fun☆13Dec 19, 2020Updated 5 years ago
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆111Updated this week
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆16Feb 21, 2026Updated 6 months ago
- Crystal and DDS Shortwave Trasmitter☆19Mar 13, 2021Updated 5 years ago
- multilingual RAG☆15Feb 6, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- 代理Groq的API服务☆13May 25, 2024Updated 2 years ago
- ☆77Feb 19, 2024Updated 2 years ago
- A proof of concept tool for using local LLMs to transform messy text documents into structured JSON☆26Aug 22, 2024Updated 2 years ago
- 引流外链【开心版引流宝】致力于为个人、团队提供基于微信私域流量的推广、引流的效率工具。可减轻人力,有效降低资源损失、流量流失的几率。引流宝完全开源,免费,可商用、可任意二次开发。引流宝可以辅助你更好地开展营销活动推广!降低运营成本,提高工作效率,获取更多资源。☆13Feb 12, 2024Updated 2 years ago
- The complete NUMA-optimized branch of the ktransformers project☆26Nov 3, 2025Updated 9 months ago
- ☆23Nov 26, 2025Updated 9 months ago
- Revision of official yolov7-pose to support custom dataset for keypoint detection☆11Nov 12, 2023Updated 2 years ago
- ☆16Oct 11, 2025Updated 10 months ago
- A simple, cross-platform CLI tool for quickly switching between Claude Code configuration profiles by managing different settings.json ve…☆17Jan 17, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 原 XiunoBBS 自带的插件☆21Apr 10, 2022Updated 4 years ago
- Enable true multi gpu capability in Comfy UI using XDiT XFuser and FSDP managed by Ray☆411Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,086Updated this week
- bge-large-zh api service☆24Jan 5, 2024Updated 2 years ago
- ☆20Aug 15, 2025Updated last year
- ChatAdmin是ChatFlow和ChatStudio的后端API服务☆18Oct 18, 2024Updated last year
- A high-performance, lightweight Go proxy that bridges the **Trae API** with the **OpenAI Standard API**. Seamlessly integrate Trae's AI c…☆17Jan 22, 2026Updated 7 months ago
- The official pytorch implementation of “Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization”.☆19May 22, 2025Updated last year
- ☆18Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- WPS Office 全家桶(文字/表格/演示/PDF)AI 助手:挂 OpenAI 兼容 / Anthropic / Codex 多家模型,AI 通过工具调用直接读写文档;支持生图、MCP、中英双语,覆盖 Windows / macOS / Linux(含国产系统)。☆44Aug 11, 2026Updated 2 weeks ago
- NeurIPS 2025☆17Sep 24, 2025Updated 11 months ago
- Workflow-to-APP、ScreenShare&FloatingVideo、GPT & 3D、SpeechRecognition&TTS☆39May 11, 2026Updated 3 months ago
- ☆58Aug 19, 2025Updated last year
- Multi-GPU device selection for LTXV2 video generation in ComfyUI☆34Jan 10, 2026Updated 7 months ago
- ☆19Mar 21, 2026Updated 5 months ago
- Implementation of SoundtStream from the paper: "SoundStream: An End-to-End Neural Audio Codec"☆13Jan 27, 2025Updated last year