π Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
β800Aug 9, 2026Updated this week
Alternatives and similar repositories for Automodel
Users that are interested in Automodel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Training library for Megatron-based models with bidirectional Hugging Face conversion capabilityβ852Updated this week
- Scalable toolkit for efficient model reinforcementβ1,891Updated this week
- A library for exporting models including NeMo and Hugging Face to optimized inference backends, and deploying them for efficient queryingβ41Updated this week
- Scalable data pre processing and curation toolkit for LLMsβ1,706Updated this week
- State-of-the-art framework for fast, large-scale training and inference of diffusion modelsβ56May 20, 2026Updated 2 months ago
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zooβ2,135Updated this week
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.β1,939Updated this week
- Accelerating MoE with IO and Tile-aware Optimizationsβ737Jul 4, 2026Updated last month
- A tool to configure, launch and manage your machine learning experiments.β251Updated this week
- β253Updated this week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,484Updated this week
- β892Updated this week
- Best practices for training DeepSeek, Mixtral, Qwen and other MoE models using Megatron Core.β202May 29, 2026Updated 2 months ago
- Bridge Megatron-Core to Hugging Face/Reinforcement Learningβ228Jun 15, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Megatron's multi-modal data loaderβ377Updated this week
- slime is an LLM post-training framework for RL Scaling.β7,820Updated this week
- A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Trainingβ903Updated this week
- Evaluate and improve models and agents using environmentsβ1,095Updated this week
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.β1,058Updated this week
- π Efficient implementations for emerging model architecturesβ5,528Updated this week
- β37Aug 7, 2025Updated last year
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculativeβ¦β3,409Updated this week
- Checkpoint-engine is a simple middleware to update model weights in LLM inference enginesβ993Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A PyTorch native platform for training generative AI modelsβ5,604Updated this week
- Distributed Compiler based on Triton for Parallel Systemsβ1,512Updated this week
- CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.β535Updated this week
- A project to improve skills of large language modelsβ1,019Updated this week
- Agentic RL on Any Harness at Scaleβ759Updated this week
- (best/better) practices of megatron on veRL and tuning guideβ139May 12, 2026Updated 2 months ago
- Minimalistic large language model 3D-parallelism trainingβ2,779May 26, 2026Updated 2 months ago
- Open-source library for scalable, reproducible evaluation of AI models and benchmarks.β322Updated this week
- A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modularβ¦β582May 18, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Ongoing research training transformer models at scaleβ17,378Updated this week
- SkyRL: A Modular Full-stack RL Library for LLMsβ2,138Updated this week
- high-performance linear attention kernel library built on TileLangβ627Updated this week
- An efficient implementation of the NSA (Native Sparse Attention) kernelβ134Jun 24, 2025Updated last year
- PyTorch Single Controllerβ1,067Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,167Updated this week
- π₯ A minimal training framework for scaling FLA modelsβ409Apr 22, 2026Updated 3 months ago