Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V
☆196Jun 30, 2026Updated 2 months ago
Alternatives and similar repositories for flash-attention-v100
Users that are interested in flash-attention-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- forked from vllm-project/flash-attention☆64May 9, 2026Updated 3 months ago
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆718Updated this week
- vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support☆22Jun 20, 2026Updated 2 months ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 5 months ago
- ☆78Feb 19, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆25Jan 30, 2026Updated 7 months ago
- Flash Attention 2 implementation for Turing GPUs☆125Mar 23, 2026Updated 5 months ago
- This suite of nodes unlocks high-performance parallel processing in ComfyUI by utilizing **Model Replication**. Unlike standard offloadin…☆62Aug 13, 2026Updated 2 weeks ago
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆115Updated this week
- Enable true multi gpu capability in Comfy UI using XDiT XFuser and FSDP managed by Ray☆418Aug 23, 2026Updated last week
- Wenzhou-Kean University AI-LAB☆10Jun 6, 2022Updated 4 years ago
- ☆18Updated this week
- Official Implementation of "GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution"☆20Apr 3, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Force any Model, VAE, or CLIP to any GPU or CPU—and keep it there (or don't).☆35Feb 6, 2026Updated 6 months ago
- ☆20May 9, 2026Updated 3 months ago
- WanImageToVideo ComfyUI node, with Tiled VAE☆16Oct 22, 2025Updated 10 months ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-reques…☆742Updated this week
- A plugin code to have a simple face instead of the Xiaozhi AI firmware☆31Aug 1, 2026Updated 3 weeks ago
- ☆155Jun 13, 2026Updated 2 months ago
- ComfyUI-Bagel is now available in ComfyUI, BAGEL is an open‑source multimodal foundation model with 7B active parameters (14B total) trai…☆30May 28, 2025Updated last year
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- extension for VS2015 that adds mixins to C#☆11Jan 8, 2017Updated 9 years ago
- Processing of astronomical images☆15Jul 17, 2026Updated last month
- Voxel-based Network for Shape Completion by Leveraging Edge Generation (ICCV 2021, oral)☆17Dec 4, 2022Updated 3 years ago
- A Rust implementation of Yolo for object detection and tracking.☆10Nov 17, 2022Updated 3 years ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Aug 24, 2026Updated last week
- 一个基于利萨茹曲线的ComfyUI自定义节点,用于模拟平滑的手持相机抖动。A ComfyUI custom node for smooth handheld camera shake simulation based on Lissajous curves.☆49Feb 28, 2026Updated 6 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆985Updated this week
- This repository is created to provide a straight forward reference for running Custom Yolov5 models on the NPU for Orangepi 5 boards - RK…☆13Feb 19, 2025Updated last year
- OpenCode plugins that help AI agents understand their identity across multi-agent sessions☆31Apr 20, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆174Updated this week
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Jul 14, 2026Updated last month
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆25Aug 22, 2026Updated last week
- GPT2# is a zero dependency, sub 1 000 loc implementation of GPT2 inference, batteries included☆13May 2, 2023Updated 3 years ago
- ☆70Nov 11, 2025Updated 9 months ago
- CSharp BPE Encoder Decoder for GPT-3☆12Dec 11, 2022Updated 3 years ago
- A C# implementation of GPT☆20Sep 14, 2024Updated last year