☆19Apr 16, 2025Updated last year
Alternatives and similar repositories for default-moe
Users that are interested in default-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 9 months ago
- An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.☆21Nov 28, 2022Updated 3 years ago
- Fork of Flame repo for training of some new stuff in development☆20Jul 15, 2026Updated last week
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Project showing how to develop NKI kernels for Llama 3.2 1B inference☆21May 29, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspace☆19Oct 21, 2024Updated last year
- [ICLR 2025] Monet: Mixture of Monosemantic Experts for Transformers☆79Jun 23, 2025Updated last year
- MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following☆16Oct 31, 2024Updated last year
- [ICML 2024] DPZero: Private Fine-Tuning of Language Models without Backpropagation☆17Sep 4, 2024Updated last year
- Official code for Cumulative Spatial Knowledge Distillation for Vision Transformers (ICCV-2023) https://openaccess.thecvf.com/content/ICC…☆15Nov 5, 2023Updated 2 years ago
- Japanese LLaMa experiment☆54Dec 27, 2025Updated 6 months ago
- Codebase for EMNLP 2025 Findings paper "Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs"☆19Nov 14, 2025Updated 8 months ago
- ☆29May 24, 2024Updated 2 years ago
- Official implementation of RMoE (Layerwise Recurrent Router for Mixture-of-Experts)☆33Aug 4, 2024Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models☆18Dec 6, 2023Updated 2 years ago
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.☆118Dec 20, 2024Updated last year
- ☆23Jan 5, 2025Updated last year
- The open-source materials for paper "Sparsing Law: Towards Large Language Models with Greater Activation Sparsity".☆32Nov 12, 2024Updated last year
- ☆14Jul 13, 2025Updated last year
- some mixture of experts architecture implementations☆27Mar 22, 2024Updated 2 years ago
- The official repository of NeurIPS'25 paper "Ada-R1: From Long-Cot to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization"☆24May 6, 2026Updated 2 months ago
- Official implementation of paper "Learning to Optimize Multi-objective Alignment Through Dynamic Reward Weighting"☆28Dec 31, 2025Updated 6 months ago
- repository for CharacterChat, a personalized social support system☆75Jul 13, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Demo WebApp using Kaldi DNN engine to convert speech to text☆11Jun 12, 2016Updated 10 years ago
- Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models☆16Feb 17, 2026Updated 5 months ago
- A summarizer for Japanese articles (but ChatGPT is better)☆10Aug 1, 2022Updated 3 years ago
- [ICML 2025] Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts☆34Nov 10, 2025Updated 8 months ago
- UnrealEngine5版VOICEVOX Engine☆13Nov 29, 2025Updated 7 months ago
- Official Codebase of "A Closer Look at Weakly-Supervised Audio-Visual Source Localization" (NeurIPS 2022)☆21Dec 6, 2022Updated 3 years ago
- Explore training for quantized models☆26Jul 12, 2025Updated last year
- Official implementation of ICLR 2025 'LORO: Parameter and Memory Efficient Pretraining via Low-rank Riemannian Optimization'☆17Apr 24, 2025Updated last year
- Repository companioning the paper "Learning Unmasking Policies for Diffusion Language Models"☆17Mar 30, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ok-Topk is a scheme for distributed training with sparse gradients. Ok-Topk integrates a novel sparse allreduce algorithm (less than 6k c…☆27Dec 10, 2022Updated 3 years ago
- ☆93Aug 18, 2024Updated last year
- ☆13Jan 27, 2019Updated 7 years ago
- Code for LAVSS: Location-Guided Audio-Visual Spatial Audio Separation☆19Feb 25, 2025Updated last year
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 5 months ago
- ☆63Jul 21, 2024Updated 2 years ago
- Official implementation of "Breaking the Factorization Barrier in Diffusion Language Models"☆17Mar 27, 2026Updated 3 months ago