☆1,583Mar 25, 2026Updated 4 months ago
Alternatives and similar repositories for megablocks
Users that are interested in megablocks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A family of open-sourced Mixture-of-Experts (MoE) Large Language Models☆1,691Mar 8, 2024Updated 2 years ago
- Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4☆1,006Jul 21, 2026Updated last week
- ☆113Aug 26, 2024Updated last year
- ☆866Dec 8, 2023Updated 2 years ago
- Minimalistic large language model 3D-parallelism training☆2,771May 26, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A PyTorch native platform for training generative AI models☆5,581Updated this week
- LLM training code for Databricks foundation models☆4,433Mar 25, 2026Updated 4 months ago
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on H…☆3,468Updated this week
- Ongoing research training transformer models at scale☆17,312Updated this week
- Tile primitives for speedy kernels☆3,588Jul 13, 2026Updated 3 weeks ago
- Ring attention implementation with flash attention☆1,042Sep 10, 2025Updated 10 months ago
- Triton-based implementation of Sparse Mixture of Experts.☆281Oct 3, 2025Updated 10 months ago
- Efficient Triton Kernels for LLM Training☆6,543Updated this week
- Fast and memory-efficient exact attention