[NSDI25] AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN Training
☆35May 2, 2025Updated last year
Alternatives and similar repositories for autoccl
Users that are interested in autoccl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A minimum demo for PyTorch distributed extension functionality for collectives.☆15Jul 29, 2024Updated 2 years ago
- ☆25Feb 12, 2025Updated last year
- [TBD] "m4: A Learned Flow-level Network Simulator" by Chenning Li, Anton A. Zabreyko, Om Chabra, Arash Nasr-Esfahany, Kevin Zhao, Pratees…☆21Jun 19, 2026Updated last month
- ☆15Oct 2, 2025Updated 10 months ago
- [ACM SIGCOMM 2024] "m3: Accurate Flow-Level Performance Estimation using Machine Learning" by Chenning Li, Arash Nasr-Esfahany, Kevin Zha…☆25Oct 2, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Venus Collective Communication Library, supported by SII and Infrawaves.☆151Jun 24, 2026Updated last month
- AI Cluster Observability & Troubleshooting Toolkit. Powered by SII & Infrawaves.☆38Apr 29, 2026Updated 3 months ago
- COCCL: Compression and precision co-aware collective communication library☆38Jul 20, 2026Updated 3 weeks ago
- Official implementation of CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training (ATC '25), built on top of Megatro…☆17Jul 6, 2025Updated last year
- TACOS: [T]opology-[A]ware [Co]llective Algorithm [S]ynthesizer for Distributed Machine Learning☆37Jun 13, 2025Updated last year
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- ONCache: A Cache-Based Low-Overhead Container Overlay Network☆21Jun 7, 2025Updated last year
- This is the repository for Direct Telemetry Access, a high-speed network telemetry collection system.☆27Apr 6, 2025Updated last year
- ☆24Sep 10, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- AI model training on heterogeneous, geo-distributed resources☆46Nov 24, 2025Updated 8 months ago
- Multipath Reliable Connection (MRC) extends InfiniBand Reliable Connection semantics so a single RDMA connection can spray traffic across…☆22Jun 8, 2026Updated 2 months ago
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design☆49Jun 27, 2026Updated last month
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,485Updated this week
- ☆1,052Apr 24, 2026Updated 3 months ago
- A SysY Compiler written by Java for the Compiler Technology Course in BUAA☆20Sep 18, 2023Updated 2 years ago
- ☆238Jul 2, 2026Updated last month
- these are custom recipes of nvidia nsight system post collection analysis.☆16Nov 7, 2025Updated 9 months ago
- RDMA and SHARP plugins for nccl library☆234Apr 3, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The source code of INFless,a native serverless platform for AI inference.☆10Oct 10, 2022Updated 3 years ago
- [NeurIPS2024] "Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design", Ruisi Cai, Yeonju Ro, Geon-Woo …☆16Dec 16, 2024Updated last year
- Atomo: Communication-efficient Learning via Atomic Sparsification☆29Dec 9, 2018Updated 7 years ago
- Efficient GPU communication over multiple NICs.☆30Nov 20, 2025Updated 8 months ago
- NS3 simulator for RDMA load balancing☆12Jan 31, 2025Updated last year
- Nex Venus Communication Library☆75Nov 17, 2025Updated 8 months ago
- ☆14Oct 23, 2023Updated 2 years ago
- Simple protocol stack based on dpdk(使用dpdk搭建协议栈)☆33Apr 25, 2024Updated 2 years ago
- htop-like TUI for real-time RDMA network monitoring.☆78Aug 4, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Explore Inter-layer Expert Affinity in MoE Model Inference☆16May 6, 2024Updated 2 years ago
- ☆67Jun 25, 2024Updated 2 years ago
- GPU-accelerated LLM Training Simulator☆52Jun 26, 2025Updated last year
- Open-source implementation for "Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow"☆94Oct 15, 2025Updated 9 months ago
- Composable and Embeddable Communication Runtime for Distributed AI Services☆102Jun 5, 2026Updated 2 months ago
- ☆56Aug 27, 2024Updated last year
- Source code of Fuyao, built on Nightcore☆17Mar 8, 2024Updated 2 years ago