PyTorch implementation of Soft MoE by Google Brain in "From Sparse to Soft Mixtures of Experts" (https://arxiv.org/pdf/2308.00951.pdf)
☆83Oct 5, 2023Updated 2 years ago
Alternatives and similar repositories for soft-mixture-of-experts
Users that are interested in soft-mixture-of-experts are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of Soft MoE, proposed by Brain's Vision team, in Pytorch☆348Apr 2, 2025Updated last year
- Implementation of ST-Moe, the latest incarnation of MoE after years of research at Brain, in Pytorch☆387Jun 17, 2024Updated 2 years ago
- Template repo for Python projects, especially those focusing on machine learning and/or deep learning.☆15Jan 14, 2026Updated 7 months ago
- ☆727Updated this week
- AdaMoLE: Adaptive Mixture of LoRA Experts☆38Oct 11, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Unofficial Implementation of Selective Attention Transformer☆20Oct 31, 2024Updated last year
- sigma-MoE layer☆21Jan 5, 2024Updated 2 years ago
- A collection of AWESOME things about mixture-of-experts☆1,291Dec 8, 2024Updated last year
- Federated Learning - PyTorch☆15Jun 27, 2021Updated 5 years ago
- ☆21Oct 31, 2022Updated 3 years ago
- Official repository for the paper "Approximating Two-Layer Feedforward Networks for Efficient Transformers"☆39Jun 11, 2025Updated last year
- [ICCV-2023] Heterogeneous Forgetting Compensation for Class-Incremental Learning☆12Dec 4, 2023Updated 2 years ago
- Federated learning using a mixture of experts☆17Feb 16, 2021Updated 5 years ago
- ☆13Jan 27, 2019Updated 7 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆25Aug 2, 2024Updated 2 years ago
- List of papers on Hallucination in LMM☆10Nov 29, 2023Updated 2 years ago
- Mixture of Attention Heads☆54Oct 10, 2022Updated 3 years ago
- Towards Understanding the Mixture-of-Experts Layer in Deep Learning☆37Dec 12, 2023Updated 2 years ago
- Pytorch-based adaptive deformable convolution☆17Jun 26, 2021Updated 5 years ago
- Reimplementation of https://github.com/montemac/algebraic_value_editing in pure PyTorch for efficiency on large models☆11Jun 28, 2023Updated 3 years ago
- GPT-2 small trained on phi-like data☆68Feb 18, 2024Updated 2 years ago
- Official code repository of Laplacian Pyramid Pansharpening Network☆26May 28, 2024Updated 2 years ago
- Provably (and non-vacuously) bounding test error of deep neural networks under distribution shift with unlabeled test data.☆10Feb 27, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Dynamic Neural Representational Decoders for High-Resolution Semantic Segmentation☆19Nov 28, 2022Updated 3 years ago
- ☆29May 4, 2024Updated 2 years ago
- Learning generative models with Sinkhorn Loss☆32Nov 9, 2018Updated 7 years ago
- Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.☆18Mar 23, 2023Updated 3 years ago
- Code repository for the paper "PanFormer: a Transformer Based Model for Pan-sharpening". ICME 2022☆29Jul 8, 2022Updated 4 years ago
- ☆14Jun 11, 2024Updated 2 years ago
- ☆18Mar 18, 2024Updated 2 years ago
- Trainable Highly-expressive Activation Functions. ECCV 2024☆40Jul 7, 2026Updated 2 months ago
- A simple but robust PyTorch implementation of RetNet from "Retentive Network: A Successor to Transformer for Large Language Models" (http…☆105Nov 24, 2023Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Machine learning project using federated learning for text generation☆11May 5, 2024Updated 2 years ago
- "Tail-Aware Sperm Analysis for Transparent Tracking of Spermatozoa" Official Implementation☆10Jan 21, 2026Updated 7 months ago
- Implementation of Switch Transformers from the paper: "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficien…☆147Aug 28, 2026Updated last week
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks☆31May 22, 2024Updated 2 years ago
- S$2$CycleDiff: Spatial-Spectral-Bilateral Cycle-Diffusion Frameworkfor Hyperspectral Image Super-Resolution(AAAI 2024)☆16Aug 14, 2025Updated last year
- DeViSE model (zero-shot learning) trained on ImageNet and deployed on AWS using Docker☆47Apr 3, 2019Updated 7 years ago
- A Pytorch implementation of Sparsely-Gated Mixture of Experts, for massively increasing the parameter count of language models☆871Sep 13, 2023Updated 2 years ago