An efficient and scalable attention module designed to reduce memory usage and improve inference speed in large language models. Designed and implemented the Multi-Head Latent Attention (MLA) module as a drop-in replacement for traditional multi-head attention (MHA) in large language models.
☆26Jun 25, 2025Updated last year
Alternatives and similar repositories for MiniGPT-and-DeepSeek-MLA-Multi-Head-Latent-Attention
Users that are interested in MiniGPT-and-DeepSeek-MLA-Multi-Head-Latent-Attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Training a BERT model from scratch.☆11Oct 15, 2023Updated 2 years ago
- Implementations of a Mixture-of-Experts (MoE) architecture designed for research on large language models (LLMs) and scalable neural netw…☆83Apr 8, 2025Updated last year
- ☆19Mar 29, 2026Updated 5 months ago
- Hierarchical Vision Transformers for Disease Progression Detection in Chest X-Ray Images☆11Jan 11, 2024Updated 2 years ago
- MMFformer: Multimodal Fusion Transformer Network for Depression Detection☆20Sep 26, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Minimal and highly hackable implementation of Looped Transformers with GPT☆25Mar 8, 2026Updated 5 months ago
- ☆12Jan 19, 2020Updated 6 years ago
- A Bigram Language Model from scratch with no-smoothing and add-one smoothing. Outputs bigram counts, bigram probabilities and probability…☆14Jan 12, 2021Updated 5 years ago
- ☆15Aug 31, 2022Updated 4 years ago
- Code for COLING 2022 paper: Modeling Intra- and Inter-Modal Relations: Hierarchical Graph Contrastive Learning for Multimodal Sentiment A…☆11May 28, 2023Updated 3 years ago
- Code of paper 《Remote Sensing Image Scene Classification Based on an Enhanced Attention Module》☆11Apr 2, 2020Updated 6 years ago
- ☆14Jan 22, 2026Updated 7 months ago
- Stanford CS144(Introduction to Computer Networking)'s project, a modern C++ implementation of TCP/IP Protocol Stack☆14Apr 2, 2024Updated 2 years ago
- Use Muon optimizer instead of AdamW.☆50Mar 2, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- minimal implementation of sft with gpt2-124M☆16Nov 10, 2025Updated 9 months ago
- Baggage Screening Scanner for dangerous objects☆14Nov 2, 2018Updated 7 years ago
- Source code of BI-Mamba for cardiovascular disease detection from two-view chest X-rays☆16Dec 10, 2025Updated 8 months ago
- Byte-Pair Encoding (BPE) (subword-based tokenization) algorithm implementaions from scratch with python☆19Jan 30, 2023Updated 3 years ago
- This is a PyTorch implementation of a Transformer Decoder based model that plays chess.☆17Mar 15, 2024Updated 2 years ago
- [Nature Machine Intelligence 2019] A deep learning approach for abnormality detection in lower extremity radiographs☆13Aug 25, 2020Updated 6 years ago
- P2PXML: Deep Geometric Framework to Predict Antibody-Antigen Binding Affinity (Journal of Structural Biology)☆15Feb 24, 2026Updated 6 months ago
- GeminiAPI_PALM_Chatbot (Audio, Vision, Prompt, Multiple-PDF and addon translator)☆14May 17, 2024Updated 2 years ago
- DDAM-PS: Diligent Domain Adaptive Mixer for Person Search -- WACV2024☆13Feb 28, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Applying PBT optimization technique to different domains☆10Oct 16, 2019Updated 6 years ago
- A simple 2D ball collision engine.☆12Jun 15, 2023Updated 3 years ago
- [ICML-2025] We introduce Lie group Relative position Encodings (LieRE) that goes beyond RoPE in supporting n-dimensional inputs.☆37Aug 13, 2025Updated last year
- Official codebase for Adaptive Online Planning for Continual Lifelong Learning.☆17Mar 26, 2020Updated 6 years ago
- ☆13Jan 11, 2024Updated 2 years ago
- The code of Adaptive Fusion Network for Remote Sensing Image Semantic Segmentation.☆11Jul 29, 2020Updated 6 years ago
- ☆24Jun 15, 2026Updated 2 months ago
- CUDA based gaussian splatting engine☆21Nov 26, 2025Updated 9 months ago
- papers about reinforcement learning☆13Jan 4, 2021Updated 5 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Implementation of BERT-based Language Models☆29Aug 3, 2026Updated 3 weeks ago
- Node based programming tool☆12Jan 8, 2023Updated 3 years ago
- Code for WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge☆17Dec 31, 2024Updated last year
- Implementation of 'Stress: Super-Resolution for Dynamic Fetal MRI using Self-Supervised Learning'☆17Nov 2, 2023Updated 2 years ago
- ☆12Jun 17, 2022Updated 4 years ago
- Official PyTorch implementation of ChAda-ViT [CVPR 2024]☆45Apr 15, 2026Updated 4 months ago
- 📸 Creates beautiful screenshots, videos, and GIFs based on terminal command output.☆17Apr 16, 2026Updated 4 months ago