Research and development for optimizing transformers
☆132Feb 16, 2021Updated 5 years ago
Alternatives and similar repositories for substation
Users that are interested in substation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DaCe - Data Centric Parallel Programming☆595Updated this week
- A Data-Centric Compiler for Machine Learning☆85Dec 14, 2025Updated 9 months ago
- An external memory allocator example for PyTorch.☆16Aug 10, 2025Updated last year
- PSTensor provides a way to hack the memory management of tensors in TensorFlow and PyTorch by defining your own C++ Tensor Class.☆10Feb 10, 2022Updated 4 years ago
- ☆79May 4, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Distributed Communication-Optimal LU-factorization Algorithm☆12Aug 1, 2021Updated 5 years ago
- PipeTransformer: Automated Elastic Pipelining for Distributed Training of Large-scale Models. ICML 2021☆56Jul 21, 2021Updated 5 years ago
- Source code repo for paper "TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation"☆10Aug 11, 2023Updated 3 years ago
- C++/MPI proxies for distributed training of deep neural networks.☆16Jun 18, 2022Updated 4 years ago
- FTPipe and related pipeline model parallelism research.☆44May 16, 2023Updated 3 years ago
- ☆13Mar 27, 2020Updated 6 years ago
- ☆19Jun 3, 2023Updated 3 years ago
- A library to analyze PyTorch traces.☆556Sep 11, 2026Updated last week
- Analyze network performance in distributed training☆20Oct 20, 2020Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆252Jul 25, 2024Updated 2 years ago
- The code for our paper "Neural Architecture Search as Program Transformation Exploration"☆17Apr 28, 2021Updated 5 years ago
- ☆13Jan 23, 2021Updated 5 years ago
- Rich editor for SDFGs with included profiling and debugging, static analysis, and interactive optimization.☆22Dec 9, 2025Updated 9 months ago
- A Chainer extension for K-FAC☆20Jun 16, 2019Updated 7 years ago
- Official repository for DistFlashAttn: Distributed Memory-efficient Attention for Long-context LLMs Training☆222Aug 19, 2024Updated 2 years ago
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications☆127May 9, 2022Updated 4 years ago
- BytePS examples (Vision, NLP, GAN, etc)☆19Nov 24, 2022Updated 3 years ago
- Standalone mini-app of the ECMWF cloud microphysics parameterization☆11Sep 2, 2026Updated 2 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Sequence-level 1F1B schedule for LLMs.☆37Aug 26, 2025Updated last year
- MONeT framework for reducing memory consumption of DNN training☆174May 4, 2021Updated 5 years ago
- A GPipe implementation in PyTorch☆866Jul 25, 2024Updated 2 years ago
- A tensor-aware point-to-point communication primitive for machine learning☆286Dec 17, 2025Updated 9 months ago
- Large scale graph learning on a single machine.☆170Feb 25, 2025Updated last year
- a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.☆1,551Jul 18, 2025Updated last year
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.☆33Apr 2, 2025Updated last year
- PyTorch extensions for high performance and large scale training.☆3,406Apr 26, 2025Updated last year
- ☆13Nov 25, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Easy integrated Python scripting embedded in C++☆24Jun 25, 2020Updated 6 years ago
- A flexible and efficient deep neural network (DNN) compiler that generates high-performance executable from a DNN model description.☆1,002Sep 19, 2024Updated 2 years ago
- paper and code for New Directions in Cloud Programming, CIDR 2021☆11Feb 17, 2021Updated 5 years ago
- ☆11Apr 29, 2023Updated 3 years ago
- optimized BERT transformer inference on NVIDIA GPU. https://arxiv.org/abs/2210.03052☆479Mar 15, 2024Updated 2 years ago
- A CPU+GPU Profiling library that provides access to timeline traces and hardware performance counters.☆992Updated this week
- Mille Crepe Bench: layer-wise performance analysis for deep learning frameworks.☆18Oct 22, 2019Updated 6 years ago