working implimention of deepseek MLA
☆44Jan 8, 2025Updated last year
Alternatives and similar repositories for Multi-Head-Latent-Attention-MLA-
Users that are interested in Multi-Head-Latent-Attention-MLA- are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Aug 19, 2024Updated 2 years ago
- ☆33Sep 22, 2025Updated 11 months ago
- A simple python package for Neural Network based on numpy☆13Sep 6, 2021Updated 5 years ago
- BigKnow2022: Bringing Language Models Up to Speed☆16Mar 27, 2023Updated 3 years ago
- This contains Matlab implementation of Johannes Kopf's image processing paper which deals with the adaptive downsampling of images. It gi…☆23Nov 1, 2017Updated 8 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Some minimal implementation of some Diffusion Models. Try to use as less code and as simple arch as possible☆21Jan 10, 2025Updated last year
- nanoGRPO is a lightweight implementation of Group Relative Policy Optimization (GRPO)☆144May 8, 2025Updated last year
- ☆17May 6, 2025Updated last year
- LoRA for convolution layer☆21Mar 9, 2023Updated 3 years ago
- Repo du cours d'introduction à l'apprentissage par renforcement.☆19Feb 2, 2025Updated last year
- An AI-powered git commit message generator written in python.☆21Feb 16, 2023Updated 3 years ago
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆16Jun 16, 2024Updated 2 years ago
- A PyTorch implementation of perceptual loss using ConvNeXt feature extractors.☆29Nov 14, 2024Updated last year
- Collection of autoregressive model implementation☆85Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆45Mar 31, 2025Updated last year
- Verilog implementation of Mersenne Twister PRNG☆32Jun 20, 2018Updated 8 years ago
- A comprehensive codebase for training and finetuning Image <> Latent models.☆49Mar 1, 2025Updated last year
- RLM (Recursive Language Model) extension for pi - process large context files that exceed LLM context windows☆19Feb 8, 2026Updated 7 months ago
- DeMo: Decoupled Momentum Optimization☆202Dec 2, 2024Updated last year
- NanoGPT-speedrunning for the poor T4 enjoyers☆72Apr 22, 2025Updated last year
- Approximating the joint distribution of language models via MCTS☆22Nov 3, 2024Updated last year
- A demo for the Direct Ascent Synthesis: Hidden Generative Capabilities in Discriminative Models paper (https://arxiv.org/abs/2502.07753)☆42Mar 5, 2025Updated last year
- An Efficent BPE Algorithm Faster then Hugging Face Tokenizer's Implementation☆13Sep 9, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆14Oct 31, 2022Updated 3 years ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 5 months ago
- Course Project for COMP4471 on RWKV☆17Feb 11, 2024Updated 2 years ago
- GSM-Symbolic templates and generated data☆92Sep 11, 2026Updated last week
- ☆25May 23, 2025Updated last year
- Complete Python bindings for the OptiX host API☆62May 12, 2026Updated 4 months ago
- The official repository of BFSR: "Boosting Flow-based Generative Super-Resolution Models via Learned Prior" [CVPR 2024]☆84Jun 13, 2024Updated 2 years ago
- A tree-based prefix cache library that allows rapid creation of looms: hierarchal branching pathways of LLM generations.☆80Feb 11, 2025Updated last year
- Hierarchical multi-agent architecture for AI agent organizations☆18Apr 9, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Fully fine-tune large models like Mistral, Llama-2-13B, or Qwen-14B completely for free☆234Oct 31, 2024Updated last year
- A graph visualization of attention☆56May 20, 2025Updated last year
- Triton Migration Guide for DeepStreamSDK.☆15Dec 19, 2023Updated 2 years ago
- A system for Prompt generation to improve Text-to-Image performance.☆99Aug 22, 2026Updated 3 weeks ago
- a KaTeX plugin for Markdown-it☆13Jun 3, 2024Updated 2 years ago
- Simple Single File Key-Value Store DB Based on SQLite☆15May 23, 2026Updated 3 months ago
- Official repo of paper LM2☆49Feb 13, 2025Updated last year