Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks (EMNLP'24)
☆143Sep 20, 2024Updated 2 years ago
Alternatives and similar repositories for Parameter-Efficient-MoE
Users that are interested in Parameter-Efficient-MoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks☆31May 22, 2024Updated 2 years ago
- 5X faster 60% less memory QLoRA finetuning☆21May 28, 2024Updated 2 years ago
- ☆279Oct 31, 2023Updated 2 years ago
- Repository for CPU Kernel Generation for LLM Inference☆28Jul 13, 2023Updated 3 years ago
- LLM-Training-API: Including Embeddings & ReRankers, mergekit, LaserRMT☆28Feb 18, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This is our own implementation of 'Layer Selective Rank Reduction'☆240May 26, 2024Updated 2 years ago
- FuseAI Project☆602Jan 25, 2025Updated last year
- ☆129Jan 22, 2024Updated 2 years ago
- [COLM 2024] LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition☆670Jul 22, 2024Updated 2 years ago
- [ACL 2024] Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models☆123May 24, 2024Updated 2 years ago
- A library for easily merging multiple LLM experts, and efficiently train the merged LLM.☆515Aug 26, 2024Updated 2 years ago
- ⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)☆1,003Dec 6, 2024Updated last year
- ☆136Aug 19, 2024Updated 2 years ago
- FuseAI Project☆93Jan 25, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Codebase for Merging Language Models (ICML 2024)☆874May 5, 2024Updated 2 years ago
- ☆17May 2, 2024Updated 2 years ago
- Tools for merging pretrained large language models.☆7,356Sep 12, 2026Updated last week
- Mixture of Expert (MoE) techniques for enhancing LLM performance through expert-driven prompt mapping and adapter combinations.☆11Feb 11, 2024Updated 2 years ago
- Load multiple LoRA modules simultaneously and automatically switch the appropriate combination of LoRA modules to generate the best answe…☆163Feb 9, 2024Updated 2 years ago
- ☆179Jul 22, 2024Updated 2 years ago
- [SIGIR'24] The official implementation code of MOELoRA.☆197Jul 22, 2024Updated 2 years ago
- Code for Zero-Shot Tokenizer Transfer☆147Jan 14, 2025Updated last year
- [ICLR 2024] Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning☆644Mar 4, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Synthetic Alphabet Dataset☆19Mar 27, 2025Updated last year
- Official PyTorch implementation of QA-LoRA☆147Mar 13, 2024Updated 2 years ago
- A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI☆770Dec 15, 2023Updated 2 years ago
- Official implementation of Half-Quadratic Quantization (HQQ)☆959Feb 26, 2026Updated 6 months ago
- ⛔ DEPRECATED -- use flash-head instead (pip install flash-head)☆29Apr 10, 2026Updated 5 months ago
- State-of-the-art Parameter-Efficient MoE Fine-tuning Method☆208Aug 22, 2024Updated 2 years ago
- A bagel, with everything.☆326Apr 11, 2024Updated 2 years ago
- ☆13Feb 18, 2024Updated 2 years ago
- [ACL 2025 🔥] Time Travel is a Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts☆19May 22, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 삼각형의 실전! Triton☆16Feb 15, 2024Updated 2 years ago
- Official repository for the paper "SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention"☆102Sep 30, 2024Updated last year
- Official code for ReLoRA from the paper Stack More Layers Differently: High-Rank Training Through Low-Rank Updates☆475Apr 21, 2024Updated 2 years ago
- Generate interleaved text and image content in a structured format you can directly pass to downstream APIs.☆29Oct 18, 2024Updated last year
- [NeurIPS 2023] LLM-Pruner: On the Structural Pruning of Large Language Models. Support Llama-3/3.1, Llama-2, LLaMA, BLOOM, Vicuna, Baich…☆1,140Oct 7, 2024Updated last year
- For releasing code related to compression methods for transformers, accompanying our publications☆460Sep 10, 2026Updated last week
- [ICML'24] Data and code for our paper "Training-Free Long-Context Scaling of Large Language Models"☆449Oct 16, 2024Updated last year