MSCCL++: A GPU-driven communication stack for scalable AI applications
☆541Jul 18, 2026Updated this week
Alternatives and similar repositories for mscclpp
Users that are interested in mscclpp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Microsoft Collective Communication Library☆394Sep 20, 2023Updated 2 years ago
- NCCL Profiling Kit☆155Jul 1, 2024Updated 2 years ago
- Microsoft Collective Communication Library☆66Nov 23, 2024Updated last year
- TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches☆83Jul 25, 2023Updated 2 years ago
- Synthesizer for optimal collective communication algorithms☆125Apr 8, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,343Aug 28, 2025Updated 10 months ago
- ☆377Jan 28, 2026Updated 5 months ago
- Distributed Compiler based on Triton for Parallel Systems☆1,493Jul 11, 2026Updated last week
- ☆47Dec 13, 2024Updated last year
- A lightweight design for computation-communication overlap.☆242Jan 20, 2026Updated 5 months ago
- Dynamic Memory Management for Serving LLMs without PagedAttention☆504Updated this week
- A GPU-driven system framework for scalable AI applications☆130Updated this week
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated 11 months ago
- Perplexity GPU Kernels☆590Nov 7, 2025Updated 8 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A throughput-oriented high-performance serving framework for LLMs☆968Mar 29, 2026Updated 3 months ago
- torchcomms: a modern PyTorch communications API☆377Updated this week
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,465Updated this week
- Unified Collective Communication Library☆310Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆419Updated this week
- ☆26Feb 17, 2025Updated last year
- NVIDIA Inference Xfer Library (NIXL)☆1,138Updated this week
- Distributed MoE in a Single Kernel [NeurIPS '25]☆272May 5, 2026Updated 2 months ago
- ☆65Apr 26, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- FlashInfer: Kernel Library for LLM Serving☆5,983Updated this week
- Thunder Research Group's Collective Communication Library☆53Jul 8, 2025Updated last year
- Nex Venus Communication Library☆75Nov 17, 2025Updated 8 months ago
- Disaggregated serving system for Large Language Models (LLMs).☆827Apr 6, 2025Updated last year
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,377Updated this week
- ☆24Feb 12, 2025Updated last year
- NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer☆195Feb 11, 2026Updated 5 months ago
- NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process com…☆560Updated this week
- Tile-based language built for AI computation across all scales☆172Jun 16, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.☆769Aug 6, 2025Updated 11 months ago
- Optimized primitives for collective multi-GPU communication☆4,892Updated this week
- A library to analyze PyTorch traces.☆535May 29, 2026Updated last month
- ☆269Jul 11, 2024Updated 2 years ago
- A low-latency & high-throughput serving engine for LLMs☆511Jan 8, 2026Updated 6 months ago
- GLake: optimizing GPU memory management and IO transmission.☆501Mar 24, 2025Updated last year
- ☆85Dec 2, 2022Updated 3 years ago