☆28Jul 11, 2021Updated 5 years ago
Alternatives and similar repositories for Optimus
Users that are interested in Optimus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Chainer extension for K-FAC☆20Jun 16, 2019Updated 7 years ago
- A Python library transfers PyTorch tensors between CPU and NVMe☆125Nov 27, 2024Updated last year
- An external memory allocator example for PyTorch.☆16Aug 10, 2025Updated last year
- Distributed Communication-Optimal Shuffle and Transpose Algorithm☆14Apr 18, 2026Updated 4 months ago
- ☆79May 4, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Release doc/tutorial/wheels for poseidon-tf☆10Jan 18, 2018Updated 8 years ago
- Distributed Communication-Optimal LU-factorization Algorithm☆12Aug 1, 2021Updated 5 years ago
- Code for "Heterogenity-Aware Cluster Scheduling Policies for Deep Learning Workloads", which appeared at OSDI 2020☆139Jul 25, 2024Updated 2 years ago
- ☆21Jul 23, 2022Updated 4 years ago
- ☆16Jan 14, 2025Updated last year
- ☆43Sep 6, 2021Updated 4 years ago
- Deep Learning ❤️ OneFlow☆19Aug 26, 2021Updated 4 years ago
- (ECCV2022) EAGAN: EAGAN: Efficient Two-stage Evolutionary Architecture Search for GANs☆12Sep 15, 2022Updated 3 years ago
- Communication patterns for AI, built on top of NCCL device and host APIs☆28Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- My solution code to parallel architecture and programming Spring 2016☆12Aug 15, 2016Updated 10 years ago
- Official repository for DistFlashAttn: Distributed Memory-efficient Attention for Long-context LLMs Training☆222Aug 19, 2024Updated last year
- Fork of cyclops-community/ctf repository updated haphazardly, previously this was main repo location☆10Aug 7, 2018Updated 8 years ago
- Scalable PaLM implementation of PyTorch☆190Dec 19, 2022Updated 3 years ago
- Complete simulation of IEEE 754 fixed and floating point specification to any precision☆13Aug 26, 2020Updated 5 years ago
- Experimental repository for a caffe2 operator☆16Dec 1, 2021Updated 4 years ago
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- PyTorch implementation of LAMB for ImageNet/ResNet-50 training☆13May 13, 2021Updated 5 years ago
- ☆12Sep 18, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Chimera: bidirectional pipeline parallelism for efficiently training large-scale models.☆72Mar 20, 2025Updated last year
- Elixir: Train a Large Language Model on a Small GPU Cluster☆16Jun 8, 2023Updated 3 years ago
- A decentralised application that creates high quality machine learning datasets☆12Jan 22, 2019Updated 7 years ago
- PatrickStar enables Larger, Faster, Greener Pretrained Models for NLP and democratizes AI for everyone.☆772Nov 18, 2025Updated 9 months ago
- Efficient Dataset Distillation by Representative Matching☆114Feb 28, 2024Updated 2 years ago
- Unofficial implementation of "A Closed-form Solution to Photorealistic Image Stylization"☆14Apr 13, 2023Updated 3 years ago
- Examples of training models with hybrid parallelism using ColossalAI☆339Mar 23, 2023Updated 3 years ago
- kubectl plugin☆14Feb 25, 2023Updated 3 years ago
- CUDA benchmarks for measuring GPU utilization and interference☆19Feb 11, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.☆33Apr 2, 2025Updated last year
- ☆39Jan 15, 2021Updated 5 years ago
- 基于FP16的二维脉动阵列电路设计☆13Feb 23, 2023Updated 3 years ago
- This aims to be an wrapper to C-MPI3 for C++, using the principles of simplicity, STL, RAII and Boost and enforcing type-safety. This i…☆23Oct 11, 2024Updated last year
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- PSTensor provides a way to hack the memory management of tensors in TensorFlow and PyTorch by defining your own C++ Tensor Class.☆10Feb 10, 2022Updated 4 years ago
- Official Pytorch code for "AesUST: Towards Aesthetic-Enhanced Universal Style Transfer" (ACM MM 2022)☆15Dec 31, 2022Updated 3 years ago