Pipeline parallelism for the minimalist
☆40Aug 6, 2025Updated last year
Alternatives and similar repositories for minPP
Users that are interested in minPP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation for the paper Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapp…☆15May 20, 2026Updated 4 months ago
- More advanced Taichi examples☆13Jun 16, 2021Updated 5 years ago
- The LLVM Symbolic Simulator, part of SAW.☆22Jul 17, 2020Updated 6 years ago
- ☆15May 24, 2025Updated last year
- [Poster; ICLR 2026] [Oral; Neurips OPT2024] μLO: Compute-Efficient Meta-Generalization of Learned Optimizers☆16Apr 15, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Sep 8, 2026Updated 2 weeks ago
- Parallel data preprocessing for NLP and ML.☆34Nov 1, 2024Updated last year
- ☆16Apr 7, 2026Updated 5 months ago
- Code release for book "Efficient Training in PyTorch"☆133Apr 10, 2025Updated last year
- Home of the Taichi documentation site.☆25Jul 17, 2026Updated 2 months ago
- PSSGen: Portable Test and Stimulus Standard DSL Generator☆15Jun 30, 2026Updated 2 months ago
- torchcomms: a modern PyTorch communications API☆398Updated this week
- SystemVerilog file list pruner☆19Sep 14, 2026Updated last week
- ☆13Apr 7, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Linter for SystemVerilog Assertions (SVA). Following the philosophy of BYOL - Build Your Own Linter, SVALint is an example of ho users ca…☆21Aug 23, 2026Updated last month
- A demo illustrating how to use Taichi as an AOT shader compiler☆78Jun 9, 2026Updated 3 months ago
- Generates a SystemVerilog assertion interface for a given SV RTL design☆20Mar 23, 2025Updated last year
- Pipeline Parallelism for PyTorch☆786Aug 21, 2024Updated 2 years ago
- This repository shows how to use Q8 kernels with `diffusers` to optimize inference of LTX-Video on ADA GPUs.☆25Jan 7, 2025Updated last year
- PyTorch bindings for CUTLASS grouped GEMM.☆155May 29, 2025Updated last year
- ☆140May 10, 2026Updated 4 months ago
- STOMP: Scheduling Techniques Optimization in heterogeneous Multi-Processors☆25Sep 17, 2025Updated last year
- This repository contains an extended version of SMCSim (originally by Erfan Azarkhish), used for near-data-processing research by Jiwon C…☆14Nov 24, 2020Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- JSON Logging for Sanic☆10Sep 1, 2021Updated 5 years ago
- Automatically insert nvtx ranges to PyTorch models☆22Apr 28, 2021Updated 5 years ago
- A codebase for pretraining multi-billion-scale sparse GPTs.☆30Feb 9, 2026Updated 7 months ago
- Memory Replay with Data Compression (ICLR 2022)☆17Sep 26, 2023Updated 2 years ago
- Chimera: bidirectional pipeline parallelism for efficiently training large-scale models.☆72Mar 20, 2025Updated last year
- ☆24Jun 18, 2024Updated 2 years ago
- [Experimental] Taichi AOT Plugin for UE5☆15Mar 15, 2023Updated 3 years ago
- Microsoft Collective Communication Library☆67Nov 23, 2024Updated last year
- (NeurIPS 2022) Automatically finding good model-parallel strategies, especially for complex models and clusters.☆45Nov 4, 2022Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICML 2026] Official implementation of "FOCUS: DLLMs Know How to Tame Their Compute Bound".☆19Aug 2, 2026Updated last month
- Compiler for Dynamic Neural Networks☆45Nov 13, 2023Updated 2 years ago
- A curated collection of resources, tutorials, and best practices for learning and mastering NVIDIA CUTLASS☆272May 6, 2025Updated last year
- Allow torch tensor memory to be released and resumed later☆276Updated this week
- Taichi-based Differentiable DVR Renderer☆12Jul 20, 2022Updated 4 years ago
- TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches☆85Jul 25, 2023Updated 3 years ago
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆336Updated this week