DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers
☆485Sep 30, 2026Updated last week
Alternatives and similar repositories for DeepSelect
Users that are interested in DeepSelect are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆28Aug 22, 2026Updated last month
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,273Sep 1, 2026Updated last month
- Compositional Muon release☆25Jun 5, 2026Updated 4 months ago
- ☆37Aug 7, 2025Updated last year
- [ICML 2026 Spotlight] Official implementation of TetraJet-v2: Accurate NVFP4 Training for LLMs, with fully-NVFP4 linear layer with unbias…☆18Jul 3, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- DeeperGEMM: crazy optimized version☆85May 5, 2025Updated last year
- ☆17May 27, 2026Updated 4 months ago
- Accepted to MLSys 2026☆97Apr 19, 2026Updated 5 months ago
- htop-like TUI for real-time RDMA network monitoring.☆113Oct 2, 2026Updated last week
- Autonomous CUDA kernel optimization agent with structured task specs and per-config scoring☆17Jun 17, 2026Updated 3 months ago
- DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling☆44Updated this week
- ☆15Feb 23, 2025Updated last year
- ☆17Oct 5, 2025Updated last year
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆211Mar 29, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- cuVS - a library for vector search and clustering on the GPU. The IVF RaBitQ is under the cuvs_ivf_rabitq branch.☆22Sep 3, 2026Updated last month
- Skills for writing tilelang and debugging with CUDA toolkits.☆145Sep 15, 2026Updated 3 weeks ago
- Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language☆369Sep 15, 2026Updated 3 weeks ago
- High-performance LLM operator library built on TileLang.☆192Updated this week
- Official repository Flash Local Linear Attention☆40May 28, 2026Updated 4 months ago
- ☆13Jan 7, 2025Updated last year
- High Performance KV Cache Store for LLM☆58May 20, 2026Updated 4 months ago
- ☆429Updated this week
- Nex Venus Communication Library☆77Nov 17, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FlashInfer Bench @ MLSys 2026: Building AI agents to write high performance GPU kernels☆184Apr 26, 2026Updated 5 months ago
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- [NeurIPS'25 Spotlight] Adaptive Attention Sparsity with Hierarchical Top-p Pruning☆110Aug 31, 2026Updated last month
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- [ICML2026] Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization☆67Jul 26, 2026Updated 2 months ago
- Cute layout visualization☆45Jan 18, 2026Updated 8 months ago
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 3 months ago
- ☆65Apr 26, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- mKernel: fast multi-node, multi-GPU fused kernels☆291Updated this week
- ☆98Feb 5, 2026Updated 8 months ago
- Awesome code, projects, books, etc. related to CUDA☆45Aug 9, 2026Updated 2 months ago
- Learning TileLang with 10 puzzles!☆386Aug 11, 2026Updated last month
- ☆54May 19, 2025Updated last year
- Distributed MoE in a Single Kernel [NeurIPS '25]☆310May 5, 2026Updated 5 months ago
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Sep 11, 2026Updated 3 weeks ago