π Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
β885Aug 30, 2026Updated this week
Alternatives and similar repositories for Automodel
Users that are interested in Automodel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Training library for Megatron-based models with bidirectional Hugging Face conversion capabilityβ890Updated this week
- Scalable toolkit for efficient model reinforcementβ1,969Updated this week
- A library for exporting models including NeMo and Hugging Face to optimized inference backends, and deploying them for efficient queryingβ42Updated this week
- Scalable data pre processing and curation toolkit for LLMsβ1,740Updated this week
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zooβ2,180Updated this week
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.β2,288Updated this week
- Accelerating MoE with IO and Tile-aware Optimizationsβ752Updated this week
- A tool to configure, launch and manage your machine learning experiments.β254Updated this week
- β269Updated this week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,511Updated this week
- An agentic-first RL framework for research (9k lines).β987Updated this week
- Best practices for training DeepSeek, Mixtral, Qwen and other MoE models using Megatron Core.β202May 29, 2026Updated 3 months ago
- Bridge Megatron-Core to Hugging Face/Reinforcement Learningβ230Jun 15, 2026Updated 2 months ago
- Megatron's multi-modal data loaderβ377Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- slime is an LLM post-training framework for RL Scaling.β8,313Updated this week
- A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Trainingβ928Updated this week
- Evaluate and improve models and agents using environmentsβ1,148Updated this week
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.β1,131Updated this week
- π Efficient implementations for emerging model architecturesβ5,665Updated this week
- β37Aug 7, 2025Updated last year
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculativeβ¦β3,612Updated this week
- Checkpoint-engine is a simple middleware to update model weights in LLM inference enginesβ1,003Aug 12, 2026Updated 2 weeks ago
- A PyTorch native platform for training generative AI modelsβ5,680Updated this week
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Distributed Compiler and Optimized Parallel Kernelsβ1,529Aug 12, 2026Updated 2 weeks ago
- CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.β543Updated this week
- A project to improve skills of large language modelsβ1,033Updated this week
- Agentic RL on Any Harness at Scaleβ819Aug 13, 2026Updated 2 weeks ago
- (best/better) practices of megatron on veRL and tuning guideβ138May 12, 2026Updated 3 months ago
- Minimalistic large language model 3D-parallelism trainingβ2,805May 26, 2026Updated 3 months ago
- Open-source library for scalable, reproducible evaluation of AI models and benchmarks.β334Updated this week
- A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modularβ¦β586May 18, 2026Updated 3 months ago
- Ongoing research training transformer models at scaleβ17,674Updated this week
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- SkyRL: A Modular Full-stack RL Library for LLMsβ2,210Updated this week
- high-performance linear attention kernel library built on TileLangβ672Updated this week
- An efficient implementation of the NSA (Native Sparse Attention) kernelβ134Jun 24, 2025Updated last year
- PyTorch Single Controllerβ1,071Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,306Updated this week
- π₯ A minimal training framework for scaling FLA modelsβ413Apr 22, 2026Updated 4 months ago
- PyTorch-native post-training at scaleβ701Updated this week