This is a simple torch implementation of the high performance Multi-Query Attention
☆16Aug 23, 2023Updated 2 years ago
Alternatives and similar repositories for MultiQueryAttention
Users that are interested in MultiQueryAttention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20Feb 2, 2026Updated 6 months ago
- [NeurIPS 2025] Encoder-Decoder Diffusion Language Models for Efficient Training and Inference☆47Oct 29, 2025Updated 9 months ago
- Demo of the unit_scaling library, showing how a model can be easily adapted to train in FP8.☆46Jul 17, 2024Updated 2 years ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization☆19Mar 7, 2025Updated last year
- The official implementation of ICLR 2025 paper "Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models".☆18Apr 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16May 18, 2026Updated 2 months ago
- ☆26May 24, 2023Updated 3 years ago
- Calculating FLOPs of Pre-trained Models in NLP☆18Mar 29, 2021Updated 5 years ago
- Code for the paper Xiangqi-R1: Enhancing Spatial Strategic Reasoning in LLMs for Chinese Chess via Reinforcement Learning☆15Jul 23, 2025Updated last year
- ☆20Oct 25, 2022Updated 3 years ago
- Code for "Multi-Objective GFlowNets"☆20Jul 12, 2023Updated 3 years ago
- ☆12Jul 25, 2020Updated 6 years ago
- Image Artisan XL is the ultimate desktop application for creating amazing images with the power of artificial intelligence.☆18Apr 25, 2024Updated 2 years ago
- ☆40May 20, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Experiments on Multi-Head Latent Attention☆101Aug 19, 2024Updated last year
- ☆12Aug 28, 2025Updated 11 months ago
- A swarm of LLM agents that will help you test, document, and productionize your code!☆20Aug 3, 2026Updated last week
- ☆24Mar 7, 2025Updated last year
- Porting Postgres Server to WASM [WIP]☆16Mar 6, 2021Updated 5 years ago
- LLM Safeguarding with Internal Representations☆20Apr 27, 2026Updated 3 months ago
- ☆24Feb 16, 2022Updated 4 years ago
- [ICLR'25] "Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers"☆44Mar 31, 2025Updated last year
- Repository for the DPP'23 course☆11May 2, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆30May 24, 2025Updated last year
- ☆21Mar 29, 2025Updated last year
- POSTECH: Compiler Construction (Spring 2022)☆11Mar 10, 2023Updated 3 years ago
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- Official implementation of the paper "You Do Not Fully Utilize Transformer's Representation Capacity"☆32May 28, 2025Updated last year
- This repository is the official implementation of Generalized Data Weighting via Class-level Gradient Manipulation (NeurIPS 2021)(http://…☆23Oct 8, 2022Updated 3 years ago
- ☆13May 11, 2023Updated 3 years ago
- Normalize CJK characters in text☆14Sep 30, 2025Updated 10 months ago
- 学习的A星算法教程,把代码分享给更多人。一起学习。☆16Apr 5, 2018Updated 8 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Explore how Flux Dev responds when you change the strengths of layers in the model.☆21Sep 20, 2024Updated last year
- Benchmark for Biophysical Sequence Optimization Algorithms☆24Apr 15, 2026Updated 3 months ago
- Implementation of SelfExtend from the paper "LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning" from Pytorch and Zeta☆13Nov 11, 2024Updated last year
- A very simple performing matrix multiplication example for CPU / CUDA / METAL using GGML / llama.cpp☆13Jul 7, 2024Updated 2 years ago
- ☆34May 4, 2026Updated 3 months ago
- [ICLR26] Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs☆33Dec 9, 2025Updated 8 months ago
- An innovative method designed to augment the capabilities of existing video diffusion models☆22May 10, 2024Updated 2 years ago