Advancing the frontier of efficient AI
☆68Jul 10, 2026Updated last month
Alternatives and similar repositories for sparse-attention-hub
Users that are interested in sparse-attention-hub are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Sep 23, 2025Updated 10 months ago
- ☆24Dec 6, 2025Updated 8 months ago
- ☆63May 19, 2025Updated last year
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆295Updated this week
- Code for Fast-weight Product Key Memory (FwPKM)☆21Mar 18, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of Paper "System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving"☆30Apr 17, 2026Updated 4 months ago
- [AAAI26]: DS SERVE: The Largest Open Vector Store over Pretain Data; A Framework for Efficient and Scalable Neural Retrieval☆54Jan 28, 2026Updated 6 months ago
- PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]☆75Oct 2, 2025Updated 10 months ago
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆25Dec 5, 2025Updated 8 months ago
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆55Oct 18, 2024Updated last year
- How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models☆87Feb 5, 2026Updated 6 months ago
- ⛔ DEPRECATED -- use flash-head instead (pip install flash-head)☆29Apr 10, 2026Updated 4 months ago
- MinT-2M: Long-context training system for resident-prefix GRPO☆44Jul 24, 2026Updated 3 weeks ago
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆119Dec 2, 2025Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆66May 7, 2026Updated 3 months ago
- Experiments Notebook of "Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism"☆16Apr 30, 2025Updated last year
- Systematic evaluation framework that automatically rates overthinking behavior in large language models.☆102May 16, 2025Updated last year
- Simulator for comparing memory allocation policies for caches.☆20May 15, 2019Updated 7 years ago
- ☆18Jul 1, 2025Updated last year
- Gecko Architecture☆18Jan 13, 2026Updated 7 months ago
- Official Repository of VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents☆114May 3, 2026Updated 3 months ago
- build and benchmark deep research☆245Mar 28, 2026Updated 4 months ago
- LLM KV cache compression made easy☆1,172Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- An LLM inference engine, written in C++☆20Mar 30, 2026Updated 4 months ago
- Repo for "AlphaResearch: Accelerating New Algorithm Discovery with Language Models"☆58Nov 12, 2025Updated 9 months ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 8 months ago
- AI-Driven Scientific, Algorithmic, and Systems Discovery☆604Updated this week
- ☆253Nov 19, 2025Updated 8 months ago
- ☆14Aug 3, 2026Updated 2 weeks ago
- Structured Primitives for Efficient Architecture Research☆21Dec 22, 2025Updated 7 months ago
- ☆15Sep 25, 2025Updated 10 months ago
- ☆15Apr 26, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- ☆328Jul 10, 2025Updated last year
- Preview Code for Continuum Paper☆98Updated this week
- Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers☆44Jul 1, 2026Updated last month
- LM engine is a library for pretraining/finetuning LLMs☆189Aug 10, 2026Updated last week
- Fast and memory-efficient classical machine learning operators☆556Aug 4, 2026Updated 2 weeks ago
- Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement Learning☆37Jul 14, 2025Updated last year