Multi-Layer Key-Value sharing experiments on Pythia models
☆34Jun 14, 2024Updated 2 years ago
Alternatives and similar repositories for pythia-mlkv
Users that are interested in pythia-mlkv are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆47Nov 25, 2024Updated last year
- Crosslingual Reasoning through Test-Time Scaling☆21May 13, 2025Updated last year
- Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging. Arxiv, 2024.☆16Oct 28, 2024Updated last year
- ☆19Mar 25, 2025Updated last year
- Official code and data repository of MathChat: MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Inte…☆22Jun 3, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- DataRubrics, a structured framework for assessing the quality of both human- and model-generated datasets. Leveraging recent advances in …☆17Jun 6, 2025Updated last year
- Token Omission Via Attention☆131Oct 13, 2024Updated last year
- LLM-Merging: Building LLMs Efficiently through Merging☆208Sep 24, 2024Updated last year
- ☆22Dec 1, 2021Updated 4 years ago
- Cross-lingual Language Model (XLM) pretraining and Model-Agnostic Meta-Learning (MAML) for fast adaptation of deep networks☆20Mar 26, 2021Updated 5 years ago
- My final project, Snow Simulation, for Prof. Lingqi Yan's online open course games 101-Intro to Modern Computer Graphics☆12Mar 12, 2021Updated 5 years ago
- Command helper for slurm system. Act as if you are on compute node.☆16Feb 1, 2025Updated last year
- Official Code Implementation for 'A Simple Early Exiting Framework for Accelerated Sampling in Diffusion Models'☆20Jul 24, 2024Updated 2 years ago
- Works about Cucker-Smale model and its extensions. =Keywords: ODE, Runge-Kutta methods, SDE, Euler-Maruyama method, NumPy, Matplotlib☆12Feb 14, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official Repository for Paper "BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Compet…☆18Aug 28, 2024Updated last year
- ☆20Mar 12, 2025Updated last year
- Repository for "Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators"☆12Mar 25, 2025Updated last year
- ☆24Mar 7, 2025Updated last year
- [EMNLP 2024] Tree of Problems: Improving structured problem solving with compositionality☆20Mar 4, 2025Updated last year
- An in-memory compressed cache for gigabytes of data written in Go.☆19Feb 6, 2023Updated 3 years ago
- [ACL 2024] RelayAttention for Efficient Large Language Model Serving with Long System Prompts☆39Feb 29, 2024Updated 2 years ago
- See the device (CPU/GPU/ANE) and estimated cost for every layer in your CoreML model.☆25Oct 23, 2025Updated 9 months ago
- ☆50May 22, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆13Mar 9, 2024Updated 2 years ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- ☆16Jul 23, 2024Updated 2 years ago
- biorbd + casadi + variational integrator☆10Jul 8, 2026Updated 3 weeks ago
- KV cache compression for high-throughput LLM inference☆158Feb 5, 2025Updated last year
- The Benefits of a Concise Chain of Thought on Problem Solving in Large Language Models☆25Nov 25, 2024Updated last year
- ☆15Mar 12, 2024Updated 2 years ago
- Live survey of off-the-shelf language identification tools for python☆27Apr 13, 2022Updated 4 years ago
- ☆38Feb 12, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The simplest implementation of recent Sparse Attention patterns for efficient LLM inference.☆92Jul 17, 2025Updated last year
- ☆14Sep 7, 2024Updated last year
- Keyformer proposes KV Cache reduction through key tokens identification and without the need for fine-tuning☆57Mar 26, 2024Updated 2 years ago
- DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching☆24Apr 15, 2026Updated 3 months ago
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.☆16May 25, 2025Updated last year
- [ICML‘2024] "LoCoCo: Dropping In Convolutions for Long Context Compression", Ruisi Cai, Yuandong Tian, Zhangyang Wang, Beidi Chen☆17Sep 7, 2024Updated last year
- Code for paper: Optimizing Length Compression in Large Reasoning Models☆29Oct 20, 2025Updated 9 months ago