MatN23 / AdaptiveTrainingSystemView on GitHub
A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD), and DeepSpeed integration. Implements models from 70M to 300B parameters with automatic dataset processing, distributed training, and memory management.
21Jul 18, 2026Updated this week

Alternatives and similar repositories for AdaptiveTrainingSystem

Users that are interested in AdaptiveTrainingSystem are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?