[CVPR2026] VecAttention: Vector-wise Sparse Attention for Accelerating Long-Context Inference
☆22May 27, 2026Updated 4 months ago
Alternatives and similar repositories for VecAttention
Users that are interested in VecAttention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆136May 17, 2026Updated 4 months ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization☆45Aug 13, 2025Updated last year
- ☆16Feb 22, 2024Updated 2 years ago
- This is a python repo for flattening Verilog☆19Dec 19, 2025Updated 9 months ago
- ☆22Nov 3, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code repo for "CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs".☆17Sep 15, 2024Updated 2 years ago
- blogs about Coimpiler & Virtual Machine☆12Jun 15, 2025Updated last year
- ☆19Feb 18, 2025Updated last year
- AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior (ACM MM 2025 Oral)☆21Mar 5, 2026Updated 6 months ago
- Official implementation of the EMNLP23 paper: Outlier Suppression+: Accurate quantization of large language models by equivalent and opti…☆52Oct 21, 2023Updated 2 years ago
- An open platform for exploring scale-up network systems.☆23Mar 16, 2026Updated 6 months ago
- Code for "RSQ: Learning from Important Tokens Leads to Better Quantized LLMs"☆24Mar 25, 2026Updated 6 months ago
- ☆17Oct 5, 2025Updated 11 months ago
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆28Sep 22, 2026Updated last week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆21Oct 13, 2024Updated last year
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆30Jul 4, 2026Updated 2 months ago
- The code of linux kernel framebuffer driver for raspberry pi and LCD screen with detailed description☆17Jun 19, 2021Updated 5 years ago
- DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching☆25Apr 15, 2026Updated 5 months ago
- ☆22Apr 27, 2026Updated 5 months ago
- [ICML 2026] Official codebase for "Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation"☆45Aug 5, 2026Updated last month
- This is the official code for the paper "EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolati…☆20May 14, 2025Updated last year
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 4 months ago
- 一个基于AXI接口的PL端卷积加速器,可由PS端调用☆12Apr 15, 2023Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [CVPR 2026 Oral] SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching☆28Jun 5, 2026Updated 3 months ago
- Python C++ Code Manager☆16Sep 29, 2024Updated 2 years ago
- Language modeling based on Penn Treebank (RNN/LSTM, Pytorch)☆16Dec 11, 2019Updated 6 years ago
- A tool for those who want to use Vivado's batch mode more easily☆17Dec 16, 2019Updated 6 years ago
- Collection of memory microbenchmarks to investigate NVIDIA GPUs Network on Chip architectures☆17Apr 14, 2026Updated 5 months ago
- ☆16Feb 11, 2025Updated last year
- Programming and Assignment Material for ECE 695☆18Apr 23, 2021Updated 5 years ago
- [ICML2026] Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization☆67Jul 26, 2026Updated 2 months ago
- Source code of paper ''KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing''☆32Oct 24, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An efficient spiking variational autoencoder☆13Nov 13, 2023Updated 2 years ago
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMs☆17Jul 18, 2024Updated 2 years ago
- A generic code base for neural network pruning, especially for pruning at initialization.☆32Sep 3, 2022Updated 4 years ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆45Jul 2, 2026Updated 3 months ago
- [ECAI 2024] MoSt-DSA: Modeling Motion and Structural Interactions for Direct Multi-Frame Interpolation in DSA Images☆19Dec 15, 2024Updated last year
- ☆25Jul 23, 2026Updated 2 months ago
- Official code for HiLS-Attention☆151Aug 13, 2026Updated last month