PyTorch implementation of Soft MoE by Google Brain in "From Sparse to Soft Mixtures of Experts" (https://arxiv.org/pdf/2308.00951.pdf)
☆83Oct 5, 2023Updated 2 years ago
Alternatives and similar repositories for soft-mixture-of-experts
Users that are interested in soft-mixture-of-experts are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PyTorch implementation of "From Sparse to Soft Mixtures of Experts"☆72Aug 22, 2023Updated 2 years ago
- Implementation of Soft MoE, proposed by Brain's Vision team, in Pytorch☆348Apr 2, 2025Updated last year
- Implementation of ST-Moe, the latest incarnation of MoE after years of research at Brain, in Pytorch☆386Jun 17, 2024Updated 2 years ago
- ☆726Jul 2, 2026Updated 3 weeks ago
- AdaMoLE: Adaptive Mixture of LoRA Experts☆38Oct 11, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Unofficial Implementation of Selective Attention Transformer☆20Oct 31, 2024Updated last year
- A collection of AWESOME things about mixture-of-experts☆1,283Dec 8, 2024Updated last year
- ☆21Oct 31, 2022Updated 3 years ago
- PyTorch implementation of moe, which stands for mixture of experts☆54Feb 11, 2021Updated 5 years ago
- [ICCV-2023] Heterogeneous Forgetting Compensation for Class-Incremental Learning☆12Dec 4, 2023Updated 2 years ago
- ☆25Aug 2, 2024Updated last year
- [ICLR 2024 Spotlight] Social Reward: Evaluating and Enhancing Generative AI through Million-User Feedback from an Online Creative Communi…☆12Mar 29, 2024Updated 2 years ago
- Experiments to assess SPADE on different LLM pipelines.☆17Apr 7, 2024Updated 2 years ago
- ☆30Sep 28, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- GPT-J 6B inference on TensorRT with INT-8 precision☆11Apr 5, 2023Updated 3 years ago
- a fast implementation of BM25☆10Sep 15, 2022Updated 3 years ago
- The code for paper: PeFoMed: Parameter Efficient Fine-tuning on Multi-modal Large Language Models for Medical Visual Question Answering☆64Dec 21, 2025Updated 7 months ago
- Implementation of "Towards Understanding Mixture of Experts in Deep Learning", NeurIPS 2022☆10Jan 6, 2023Updated 3 years ago
- Towards Understanding the Mixture-of-Experts Layer in Deep Learning☆35Dec 12, 2023Updated 2 years ago
- Reimplementation of https://github.com/montemac/algebraic_value_editing in pure PyTorch for efficiency on large models☆11Jun 28, 2023Updated 3 years ago
- Provably (and non-vacuously) bounding test error of deep neural networks under distribution shift with unlabeled test data.☆10Feb 27, 2024Updated 2 years ago
- Included the standard NMS, Rotate NMS,standard SoftNMS, Rotate SoftNMS, Weighted NMS, Rotate Weighted NMS, Weighted-Boxes-Fusion, Rotate …☆15Oct 11, 2021Updated 4 years ago
- A curated reading list of research in Mixture-of-Experts(MoE).☆670Oct 30, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Code repository for the paper "PanFormer: a Transformer Based Model for Pan-sharpening". ICME 2022☆29Jul 8, 2022Updated 4 years ago
- ☆14Jun 11, 2024Updated 2 years ago
- ☆18Mar 18, 2024Updated 2 years ago
- "Tail-Aware Sperm Analysis for Transparent Tracking of Spermatozoa" Official Implementation☆10Jan 21, 2026Updated 6 months ago
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks☆31May 22, 2024Updated 2 years ago
- automatic analysis of scanned documents(only extracting handwritten digits)☆17Jul 24, 2021Updated 5 years ago
- DeViSE model (zero-shot learning) trained on ImageNet and deployed on AWS using Docker☆47Apr 3, 2019Updated 7 years ago
- A Pytorch implementation of Sparsely-Gated Mixture of Experts, for massively increasing the parameter count of language models☆866Sep 13, 2023Updated 2 years ago
- ☆19Jun 10, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- FUSION is an open-source project aimed at revolutionizing networking through the simulation of advanced SD-EONs and AI-enhanced networks,…☆15Jun 23, 2026Updated last month
- This repository contains the source code of our proposed multimodal image segmentation frameworks. The network architectures and training…☆22Sep 18, 2025Updated 10 months ago
- SpringBoot Starter for Twitter4J☆12Sep 18, 2025Updated 10 months ago
- ☆10Mar 4, 2024Updated 2 years ago
- Adapter-X: A Novel General Parameter-Efficient Fine-Tuning Framework for Vision☆11Jul 22, 2024Updated 2 years ago
- Tiny ResNet inspired FPN network (<2M params) for Rotated Object Detection using 5-parameter Modulated Rotation Loss☆18Jul 8, 2021Updated 5 years ago
- Implementation of the paper: "Mixture-of-Depths: Dynamically allocating compute in transformer-based language models"☆123Updated this week