☆20Apr 16, 2025Updated last year
Alternatives and similar repositories for default-moe
Users that are interested in default-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The code of Advancing Expert Specialization for Better MoE (NeurIPS2025 oral)☆36Jan 22, 2026Updated 6 months ago
- PyCUDA based PyTorch Extension Made Easy☆27Mar 22, 2024Updated 2 years ago
- ☆15Oct 4, 2024Updated last year
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 9 months ago
- Fork of Flame repo for training of some new stuff in development☆20Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Project showing how to develop NKI kernels for Llama 3.2 1B inference☆21May 29, 2025Updated last year
- Code and data release for the paper "Seeing the Arrow of Time in Large Multimodal Models"☆16Oct 2, 2025Updated 10 months ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspace☆19Oct 21, 2024Updated last year
- [ICLR 2025] Monet: Mixture of Monosemantic Experts for Transformers☆79Jun 23, 2025Updated last year
- MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following☆16Oct 31, 2024Updated last year
- maum-ai.github.io☆15Jun 12, 2026Updated 2 months ago
- ☆19Aug 4, 2025Updated last year
- Training hybrid models for dummies.☆31Nov 1, 2025Updated 9 months ago
- ☆14May 4, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official code for Cumulative Spatial Knowledge Distillation for Vision Transformers (ICCV-2023) https://openaccess.thecvf.com/content/ICC…☆15Nov 5, 2023Updated 2 years ago
- 6,080-param transformer achieving 100% accuracy on 10-digit addition. Trained from scratch in 10 minutes.☆22Feb 19, 2026Updated 5 months ago
- [CVPR 2026 Findings] MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing☆15Updated this week
- Codebase for EMNLP 2025 Findings paper "Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs"☆19Nov 14, 2025Updated 9 months ago
- [ICLR 2025] COAT: Compressing Optimizer States and Activation for Memory-Efficient FP8 Training☆263Aug 9, 2025Updated last year
- Official implementation of RMoE (Layerwise Recurrent Router for Mixture-of-Experts)☆33Aug 4, 2024Updated 2 years ago
- aigc evals☆10Dec 2, 2023Updated 2 years ago
- Pytorch Text GAN for lyrics generation☆10Apr 13, 2019Updated 7 years ago
- [NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models☆18Dec 6, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Jan 5, 2025Updated last year
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.☆121Dec 20, 2024Updated last year
- ☆19Jan 4, 2024Updated 2 years ago
- ☆13Mar 15, 2022Updated 4 years ago
- Implementation of Multiscreen proposed by Ken Nakanishi for "Screening is Enough"☆18May 13, 2026Updated 3 months ago
- SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining (NeurIPS 2024)☆39Nov 1, 2024Updated last year
- ☆14Jul 13, 2025Updated last year
- some mixture of experts architecture implementations☆28Mar 22, 2024Updated 2 years ago
- Repository for the 2023 WACV paper: "Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source Localization"☆12Dec 21, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Apr 3, 2026Updated 4 months ago
- The official repository of NeurIPS'25 paper "Ada-R1: From Long-Cot to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization"☆24May 6, 2026Updated 3 months ago
- Official implementation of paper "Learning to Optimize Multi-objective Alignment Through Dynamic Reward Weighting"☆28Dec 31, 2025Updated 7 months ago
- UnrealEngine5版VOICEVOX Engine☆13Nov 29, 2025Updated 8 months ago
- Official Codebase of "A Closer Look at Weakly-Supervised Audio-Visual Source Localization" (NeurIPS 2022)☆22Dec 6, 2022Updated 3 years ago
- Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation☆15Aug 27, 2024Updated last year
- Explore training for quantized models☆28Jul 12, 2025Updated last year