[NSDI25] AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN Training
☆32May 2, 2025Updated last year
Alternatives and similar repositories for autoccl
Users that are interested in autoccl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17May 21, 2025Updated last year
- A minimum demo for PyTorch distributed extension functionality for collectives.☆15Jul 29, 2024Updated last year
- ☆24Feb 12, 2025Updated last year
- [TBD] "m4: A Learned Flow-level Network Simulator" by Chenning Li, Anton A. Zabreyko, Om Chabra, Arash Nasr-Esfahany, Kevin Zhao, Pratees…☆21Jun 19, 2026Updated last month
- [ACM SIGCOMM 2024] "m3: Accurate Flow-Level Performance Estimation using Machine Learning" by Chenning Li, Arash Nasr-Esfahany, Kevin Zha…☆25Oct 2, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Venus Collective Communication Library, supported by SII and Infrawaves.☆151Jun 24, 2026Updated 3 weeks ago
- COCCL: Compression and precision co-aware collective communication library☆36Jul 7, 2026Updated 2 weeks ago
- Code for "Practical Low-Rank Communication Compression in Decentralized Deep Learning"☆17Aug 4, 2020Updated 5 years ago
- Official implementation of CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training (ATC '25), built on top of Megatro…☆17Jul 6, 2025Updated last year
- TACOS: [T]opology-[A]ware [Co]llective Algorithm [S]ynthesizer for Distributed Machine Learning☆37Jun 13, 2025Updated last year
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- This is the repository for Direct Telemetry Access, a high-speed network telemetry collection system.☆27Apr 6, 2025Updated last year
- ☆24Sep 10, 2025Updated 10 months ago
- AI model training on heterogeneous, geo-distributed resources☆46Nov 24, 2025Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Multipath Reliable Connection (MRC) extends InfiniBand Reliable Connection semantics so a single RDMA connection can spray traffic across…☆21Jun 8, 2026Updated last month
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design☆48Jun 27, 2026Updated 3 weeks ago
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,467Updated this week
- ☆237Jul 2, 2026Updated 2 weeks ago
- GPU-accelerated LLM Training Simulator☆52Jun 26, 2025Updated last year
- A SysY Compiler written by Java for the Compiler Technology Course in BUAA☆20Sep 18, 2023Updated 2 years ago
- these are custom recipes of nvidia nsight system post collection analysis.☆16Nov 7, 2025Updated 8 months ago
- A Sysy Compiler Judge. By @ap0stader & @swkfk☆22Jan 12, 2025Updated last year
- Managed collective communication service☆24Sep 2, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- TraceWeaver is a research prototype for transparently tracing requests through a microservice without application instrumentation.☆23Sep 2, 2024Updated last year
- [NeurIPS2024] "Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design", Ruisi Cai, Yeonju Ro, Geon-Woo …☆16Dec 16, 2024Updated last year
- Atomo: Communication-efficient Learning via Atomic Sparsification☆29Dec 9, 2018Updated 7 years ago
- Efficient GPU communication over multiple NICs.☆29Nov 20, 2025Updated 8 months ago
- [ICLR 2025] DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models☆19Mar 25, 2025Updated last year
- NS3 simulator for RDMA load balancing☆12Jan 31, 2025Updated last year
- ☆14Oct 23, 2023Updated 2 years ago
- Nex Venus Communication Library☆75Nov 17, 2025Updated 8 months ago
- Simple protocol stack based on dpdk(使用dpdk搭建协议栈)☆33Apr 25, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Explore Inter-layer Expert Affinity in MoE Model Inference☆16May 6, 2024Updated 2 years ago
- ☆67Jun 25, 2024Updated 2 years ago
- htop-like TUI for real-time RDMA network monitoring.☆76Jul 12, 2026Updated last week
- Network components (NIC, Switch) for FireBox☆19Oct 27, 2024Updated last year
- Perplexity GPU Kernels☆591Nov 7, 2025Updated 8 months ago
- Open-source implementation for "Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow"☆93Oct 15, 2025Updated 9 months ago
- Compression for Foundation Models☆36Jul 21, 2025Updated 11 months ago