Ongoing research training transformer models at scale
β396Aug 20, 2024Updated last year
Alternatives and similar repositories for Megatron-LM
Users that are interested in Megatron-LM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Home of StarCoder: fine-tuning & inference!β7,503Feb 27, 2024Updated 2 years ago
- π OctoPack: Instruction Tuning Code Large Language Modelsβ479Feb 5, 2025Updated last year
- A framework for the evaluation of autoregressive code generation language models.β1,054Jul 22, 2025Updated last year
- β497Aug 15, 2024Updated 2 years ago
- CodeGen is a family of open-source model for program synthesis. Trained on TPU-v4. Competitive with OpenAI Codex.β5,179Jun 2, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- LLMs build upon Evol Insturct: WizardLM, WizardCoder, WizardMathβ9,482Jun 7, 2025Updated last year
- CodeGen2 models for program synthesisβ268Jun 2, 2026Updated 2 months ago
- Fine-tune SantaCoder for Code/Text Generation.β196Apr 11, 2023Updated 3 years ago
- β15Oct 24, 2023Updated 2 years ago
- Repository for analysis and experiments in the BigCode project.β126Mar 20, 2024Updated 2 years ago
- Ongoing research training transformer language models at scale, including: BERT & GPT-2β1,447Mar 20, 2024Updated 2 years ago
- This is the official code for the paper CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning (Neurβ¦β575Jun 2, 2026Updated 2 months ago
- distributed trainer for LLMsβ590May 20, 2024Updated 2 years ago
- Central place for the engineering/scaling WG: documentation, SLURM scripts and logs, compute environment and data.β1,019Jul 29, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β26Mar 6, 2024Updated 2 years ago
- CodeTF: One-stop Transformer Library for State-of-the-art Code LLMβ1,480May 1, 2025Updated last year
- Dromedary: towards helpful, ethical and reliable LLMs.β1,139Sep 18, 2025Updated 11 months ago
- LLM training code for Databricks foundation modelsβ4,440Mar 25, 2026Updated 4 months ago
- APPS: Automated Programming Progress Standard (NeurIPS 2021)β537Jun 19, 2024Updated 2 years ago
- β39Oct 3, 2022Updated 3 years ago
- β1,514May 12, 2023Updated 3 years ago
- A multi-programming language benchmark for LLMsβ313Apr 12, 2026Updated 4 months ago
- Large Language Model Text Generation Inferenceβ10,886Mar 21, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,511May 1, 2026Updated 3 months ago
- Ongoing research training transformer models at scaleβ17,466Updated this week
- C++ implementation for π«StarCoderβ459Sep 9, 2023Updated 2 years ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)β4,753Jan 8, 2024Updated 2 years ago
- Implementation of the LLaMA language model based on nanoGPT. Supports flash attention, Int8 and GPTQ 4bit quantization, LoRA and LLaMA-Adβ¦β6,084Jul 1, 2025Updated last year
- Code used for sourcing and cleaning the BigScience ROOTS corpusβ318Mar 20, 2023Updated 3 years ago
- β12Oct 7, 2023Updated 2 years ago
- β287Apr 25, 2023Updated 3 years ago
- Astraios: Parameter-Efficient Instruction Tuning Code Language Modelsβ63Apr 10, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Code for the TMLR 2023 paper "PPOCoder: Execution-based Code Generation using Deep Reinforcement Learning"β116Jan 9, 2024Updated 2 years ago
- Scaling Data-Constrained Language Modelsβ346Jun 28, 2025Updated last year
- An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed librariesβ7,453Jun 11, 2026Updated 2 months ago
- LLM powered development for VSCodeβ1,313May 26, 2026Updated 2 months ago
- Home of CodeT5: Open Code LLMs for Code Understanding and Generationβ3,097Jun 25, 2026Updated last month
- Code for EMNLP'24 paper - On Diversified Preferences of Large Language Model Alignmentβ16Aug 6, 2024Updated 2 years ago
- OpenLLaMA, a permissively licensed open source reproduction of Meta AIβs LLaMA 7B trained on the RedPajama datasetβ7,530Jul 16, 2023Updated 3 years ago