PyTorch DeepSeek Sparse Attention (DSA) training & inference
☆23Oct 14, 2025Updated 9 months ago
Alternatives and similar repositories for deepseek-sparse-attention-pytorch
Users that are interested in deepseek-sparse-attention-pytorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeepSeek Native Sparse Attention pytorch implementation☆119Dec 17, 2025Updated 7 months ago
- Efficient triton implementation of Native Sparse Attention.☆283May 23, 2025Updated last year
- ICCV2023 Rethinking Fast Fourier Convolution in Image Inpainting☆36Dec 24, 2024Updated last year
- 微信小程序之记事本☆11Jun 6, 2018Updated 8 years ago
- [WIP] Better (FP8) attention for Hopper☆33Feb 24, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- DeepSeek-V3.2-Exp DSA Warmup Lightning Indexer training operator based on tilelang☆47Nov 19, 2025Updated 8 months ago
- The Python code of a deep-unfolding algorithm for weighted sum rate maximization (WSRMax) precoding design in multiuser MIMO systems☆13Apr 25, 2023Updated 3 years ago
- Hardware CD/CI and Development Containers 🚢☆11Jul 20, 2022Updated 4 years ago
- ☆16Dec 11, 2025Updated 7 months ago
- Repository for the CVPR23 paper Re^2TAL☆13Nov 21, 2025Updated 8 months ago
- ☆14Dec 14, 2022Updated 3 years ago
- 魔镜魔镜,无所不知的魔镜[-_-](并不是)☆13Jun 10, 2021Updated 5 years ago
- ☆15Jul 6, 2024Updated 2 years ago
- ☆16Mar 10, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is the official code for UGTs.☆13Feb 8, 2023Updated 3 years ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- ☆13May 30, 2024Updated 2 years ago
- Matrix multiplication on multiple Nios II cores☆16Feb 12, 2020Updated 6 years ago
- An implementation is provided here for the NeurIPS2024 paper "MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected…☆16Mar 24, 2026Updated 4 months ago
- Self Reproduction Code of Paper "Reducing Transformer Key-Value Cache Size with Cross-Layer Attention (MIT CSAIL)☆17May 24, 2024Updated 2 years ago
- ☆36Aug 23, 2023Updated 2 years ago
- Using ch_PP-OCRv4 model with tensorrt☆17Apr 16, 2026Updated 3 months ago
- ☆16Dec 16, 2021Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"☆14Jun 7, 2023Updated 3 years ago
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Jul 17, 2026Updated 3 weeks ago
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- 使用ISTANet网络重构压缩感知采样的1维光谱信号☆14Oct 12, 2023Updated 2 years ago
- Position sensitive PreciseRoIPooling without roi coordinates gradient backward☆16Aug 2, 2018Updated 8 years ago
- ☆17Feb 24, 2019Updated 7 years ago
- ☆16Jun 7, 2022Updated 4 years ago
- Package Xilinx FPGA tools into docker containers, useful for CI situations.☆17Oct 20, 2014Updated 11 years ago
- codes for FreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference Strategy (CVPR2025)☆22Jun 18, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆12Dec 8, 2016Updated 9 years ago
- Developed a high-performance triangle-to-quad conversion operator, formulated as a maximum-weight matching problem on the triangle adjace…☆17May 28, 2026Updated 2 months ago
- Hyperledger Indy/Sovrin/DID Comprehensive Architecture Reference Model (INDY ARM) - Draft document for discussion purposes☆14Jan 25, 2021Updated 5 years ago
- [NeurIPS 2025] Code for Low-Rank Head Avatar Personalization with Registers☆19Dec 9, 2025Updated 8 months ago
- Deep unfolded SCA for power allocation in wireless system☆17Mar 14, 2022Updated 4 years ago
- 一个基于多Agent架构的智能助手系统,集成了多种MCP(Model Context Protocol)工具,通过AI Agents协作完成复杂任务。☆15Dec 26, 2025Updated 7 months ago
- Memory-Augmented Deep Unfolding Network for Compressive Sensing(ACMMM 2021)☆15Apr 1, 2022Updated 4 years ago