β20Nov 5, 2024Updated last year
Alternatives and similar repositories for MH-MoE
Users that are interested in MH-MoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Modelβ13Feb 11, 2025Updated last year
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β17Feb 4, 2025Updated last year
- Official Implementation for the paper "VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models"β24Aug 14, 2025Updated last year
- β18Mar 2, 2026Updated 6 months ago
- Mixture of Lora Expertsβ11Apr 7, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.β123Dec 20, 2024Updated last year
- Official implementation of the ICLR'25 paper "QERA: an Analytical Framework for Quantization Error Reconstruction".β14Feb 4, 2025Updated last year
- β74Dec 2, 2024Updated last year
- Special Function Units (SFUs) are hardware accelerators, their implementation helps improve the performance of GPUs to process some of thβ¦β18Sep 21, 2025Updated last year
- β22Jul 30, 2024Updated 2 years ago
- Google DeepMind: Mixture of Depths Unofficial Implementation.β12May 29, 2024Updated 2 years ago
- The code of SKSβ15Mar 22, 2022Updated 4 years ago
- ControlLM is a method to control the personality traits and behaviors of language models in real-time at inference without costly traininβ¦β21Nov 6, 2024Updated last year
- β14Mar 6, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An example of DyNet autobatching for the NIPS "how to code a paper" workshopβ12Dec 9, 2017Updated 8 years ago
- Public version for DistPepFoldβ10Jul 17, 2025Updated last year
- β10Nov 17, 2020Updated 5 years ago
- β20Dec 14, 2024Updated last year
- β13Aug 20, 2021Updated 5 years ago
- β10Jun 4, 2021Updated 5 years ago
- Predicting Protein β Ligand Interaction by using Deep Learning Modelsβ11Nov 13, 2018Updated 7 years ago
- AdaMoLE: Adaptive Mixture of LoRA Expertsβ38Oct 11, 2024Updated last year
- PyTorch implementation of "From Sparse to Soft Mixtures of Experts"β72Aug 22, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- Mixture of Attention Headsβ54Oct 10, 2022Updated 3 years ago
- [ICLR 2026] Geometric-Mean Policy Optimizationβ105Jan 26, 2026Updated 8 months ago
- β21Mar 11, 2026Updated 6 months ago
- β51Jul 3, 2026Updated 2 months ago
- Converts AlphaFold distograms into distance matrices and saves them into a number of formatsβ16Dec 13, 2022Updated 3 years ago
- β18Nov 25, 2024Updated last year
- Distributed Network Architecture Searchβ10Oct 13, 2019Updated 6 years ago
- This repo is for the paper: On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmarkβ25Aug 13, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A template for running Stable Diffusion 3 with Cogβ14Aug 20, 2024Updated 2 years ago
- ProtTrans is providing state of the art pretrained language models for proteins. ProtTrans was trained on thousands of GPUs from Summit aβ¦β12Jun 2, 2022Updated 4 years ago
- hierarchical convolutional attention networks for text classificationβ16Aug 1, 2019Updated 7 years ago
- A Winograd Minimal Filter Implementation in CUDAβ31Aug 25, 2021Updated 5 years ago
- C compiler created in Python and LLVMβ20Jun 9, 2021Updated 5 years ago
- FurNet: A Deep-Learning-Based Framework for Removing Furniture Objects in Room Imageβ13Nov 22, 2022Updated 3 years ago
- The program will output the home and away teams as well as their respective score predictions.β13Oct 18, 2020Updated 5 years ago