Pipeline parallelism for the minimalist
☆39Aug 6, 2025Updated last year
Alternatives and similar repositories for minPP
Users that are interested in minPP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- More advanced Taichi examples☆13Jun 16, 2021Updated 5 years ago
- The LLVM Symbolic Simulator, part of SAW.☆22Jul 17, 2020Updated 6 years ago
- ☆15May 24, 2025Updated last year
- Generate ctypes boilerplate code from debugging information; Use python to mock C code for testing☆29Mar 25, 2025Updated last year
- ☆42Dec 10, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Jun 11, 2018Updated 8 years ago
- ☆23Jun 9, 2026Updated 2 months ago
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Jan 30, 2026Updated 6 months ago
- Code release for book "Efficient Training in PyTorch"☆133Apr 10, 2025Updated last year
- PSSGen: Portable Test and Stimulus Standard DSL Generator☆15Jun 30, 2026Updated last month
- torchcomms: a modern PyTorch communications API☆385Updated this week
- SystemVerilog file list pruner☆18Aug 2, 2026Updated last week
- Linter for SystemVerilog Assertions (SVA). Following the philosophy of BYOL - Build Your Own Linter, SVALint is an example of ho users ca…☆21Sep 10, 2025Updated 11 months ago
- This repository provides the IEEE 1685 IP-XACT schema files for a Git submodule integration.☆21May 12, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Complete simulation of IEEE 754 fixed and floating point specification to any precision☆13Aug 26, 2020Updated 5 years ago
- Generates a SystemVerilog assertion interface for a given SV RTL design☆20Mar 23, 2025Updated last year
- A codebase for pretraining multi-billion-scale sparse GPTs.☆27Feb 9, 2026Updated 6 months ago
- Pipeline Parallelism for PyTorch☆786Aug 21, 2024Updated last year
- PyTorch bindings for CUTLASS grouped GEMM.☆154May 29, 2025Updated last year
- ☆128May 10, 2026Updated 3 months ago
- Automatically insert nvtx ranges to PyTorch models☆22Apr 28, 2021Updated 5 years ago
- Chimera: bidirectional pipeline parallelism for efficiently training large-scale models.☆72Mar 20, 2025Updated last year
- [Experimental] Taichi AOT Plugin for UE5☆15Mar 15, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Microsoft Collective Communication Library☆66Nov 23, 2024Updated last year
- (NeurIPS 2022) Automatically finding good model-parallel strategies, especially for complex models and clusters.☆44Nov 4, 2022Updated 3 years ago
- [ICML 2026] Official implementation of "FOCUS: DLLMs Know How to Tame Their Compute Bound".☆18Aug 2, 2026Updated last week
- Compiler for Dynamic Neural Networks☆45Nov 13, 2023Updated 2 years ago
- A curated collection of resources, tutorials, and best practices for learning and mastering NVIDIA CUTLASS☆269May 6, 2025Updated last year
- Allow torch tensor memory to be released and resumed later☆267Updated this week
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆320Updated this week
- TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches☆83Jul 25, 2023Updated 3 years ago
- Soft2D: A 2D multi-material continuum physics engine designed for real-time applications.☆52Aug 21, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆546Updated this week
- Dr Sparsh's lecture slides on RISC-V ISA☆22Jun 9, 2025Updated last year
- STREAMer: Benchmarking remote volatile and non-volatile memory bandwidth☆18Aug 21, 2023Updated 2 years ago
- Example Achilles SDK controller for tutorial purposes.☆14Jun 18, 2025Updated last year
- ☆143Jan 30, 2025Updated last year
- random command line tools for deep learning☆10May 13, 2026Updated 2 months ago
- An Efficient Pipelined Data Parallel Approach for Training Large Model☆76Dec 11, 2020Updated 5 years ago