☆15Sep 24, 2026Updated 2 weeks ago
Alternatives and similar repositories for DeepEP
Users that are interested in DeepEP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Updated this week
- Modular RDMA Interface☆185Updated this week
- ☆77Updated this week
- ☆24Sep 23, 2026Updated 2 weeks ago
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This repository contains an implementation for Portals4. Portals4 is a Network Programming Interface which allows high-performance networ…☆14Sep 3, 2024Updated 2 years ago
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆202Updated this week
- 14 basic topics for VEGA64 performance optmization☆66Mar 18, 2021Updated 5 years ago
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆131Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆94Sep 28, 2026Updated last week
- Header-only C++20 wrapper for MPI 4.0.☆16Oct 20, 2023Updated 2 years ago
- AiTer Optimized Model☆190Updated this week
- C++/MPI proxies for distributed training of deep neural networks.☆16Jun 18, 2022Updated 4 years ago
- A minimum demo for PyTorch distributed extension functionality for collectives.☆15Jul 29, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- CacheDirector - Sending Packets to the Right Slice by Exploiting Intel Last-Level Cache Addressing☆11Apr 29, 2019Updated 7 years ago
- Hands-on HPC I/O tutorial material☆19Aug 6, 2026Updated 2 months ago
- FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.☆18Nov 20, 2024Updated last year
- Sources and examples for ASPLOS20 paper☆14Jul 21, 2020Updated 6 years ago
- OFED libibverbs tests package☆17Oct 5, 2021Updated 5 years ago
- ☆15Apr 18, 2023Updated 3 years ago
- OCaml/MPI interface☆29Mar 8, 2026Updated 7 months ago
- A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch☆24Sep 30, 2026Updated last week
- Multi-Agent Optimization in Python☆22Nov 29, 2019Updated 6 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆17May 8, 2021Updated 5 years ago
- Compact LSTM inference kernel (CLINK) designed in C/HLS for FPGA implementation.☆17Oct 7, 2019Updated 7 years ago
- ☆21Jul 5, 2024Updated 2 years ago
- NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process com…☆596Updated this week
- Open source version of DOCA GPUNetIO and DOCA Verbs libraries (limited features) to enable GDAKI technology on RDMA (IB and RoCE)☆76Updated this week
- Not regularly updated clone of http://git.dpdk.org/dpdk-stable/ with the purpose to develop a new driver for corundum/mqnic (https://gith…☆16Aug 24, 2023Updated 3 years ago
- plget is a tool used to measure latency packets spent in network stack, NIC driver and on the wire, trace interpacket gap, based as on h/…☆19Nov 18, 2019Updated 6 years ago
- Trained a 114 million Parameter LLM from Scratch.☆19Jul 21, 2024Updated 2 years ago
- AI Tensor Engine for ROCm☆586Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [EuroSys'25] Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization☆24Apr 13, 2026Updated 5 months ago
- ☆17Oct 17, 2025Updated 11 months ago
- ☆17Mar 17, 2023Updated 3 years ago
- Header-only plugin for the Google Test framework defining listener(s) emitting sensible output when testing MPI-based, distributed-memory…☆23Jun 12, 2021Updated 5 years ago
- Enhanced PQOS (Intel RDT Software) with DDIO-related Functionalities☆16May 25, 2022Updated 4 years ago
- Course project for Operating Systems at XJTU: A basic x86-64 dynamic linker.☆13May 10, 2023Updated 3 years ago
- fake CUTLASS to get peformance☆25Apr 28, 2026Updated 5 months ago