Manages vllm-nccl dependency
☆19Jun 3, 2024Updated 2 years ago
Alternatives and similar repositories for vllm-nccl
Users that are interested in vllm-nccl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PSTensor provides a way to hack the memory management of tensors in TensorFlow and PyTorch by defining your own C++ Tensor Class.☆10Feb 10, 2022Updated 4 years ago
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Jul 17, 2026Updated last month
- An auxiliary project analysis of the characteristics of KV in DiT Attention.☆34Nov 29, 2024Updated last year
- ☆11Apr 3, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆20May 14, 2025Updated last year
- Elixir: Train a Large Language Model on a Small GPU Cluster☆16Jun 8, 2023Updated 3 years ago
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆13Apr 17, 2025Updated last year
- A Triton-only attention backend for vLLM☆28Jul 14, 2026Updated last month
- ☆21Mar 22, 2021Updated 5 years ago
- A Feishu/Lark AI agent bot☆15Feb 27, 2026Updated 5 months ago
- ☆13Jul 24, 2024Updated 2 years ago
- ☆47Dec 13, 2024Updated last year
- Depict GPU memory footprint during DNN training of PyTorch☆11Nov 17, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- SmartNIC☆14Dec 13, 2018Updated 7 years ago
- A Suite for Parallel Inference of Diffusion Transformers (DiTs) on multi-GPU Clusters☆58May 3, 2026Updated 3 months ago
- GPT Demo with hybrid distributed training☆10Dec 1, 2022Updated 3 years ago
- ☆15Dec 1, 2023Updated 2 years ago
- PyTorch compilation tutorial covering TorchScript, torch.fx, and Slapo☆17Mar 13, 2023Updated 3 years ago
- Official repo for "Binary Retrieval-augmented Reward Mitigates Hallucinations"☆16Nov 13, 2025Updated 9 months ago
- Debug DeepSpeed-Chat step by step in IDE (在IDE里一步一步调试DeepSpeed-Chat)☆10Apr 17, 2023Updated 3 years ago
- ☆20Sep 28, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆28Jul 29, 2025Updated last year
- This is an official GitHub repository for the paper, "Towards timeout-less transport in commodity datacenter networks.".☆17Oct 12, 2021Updated 4 years ago
- 3位代码类目表;6位扩展代码表;疾病分类与代码(修订版);章节名称及代码☆11Aug 20, 2018Updated 7 years ago
- Drop-in library for tracking the memory allocations of CUDA applications☆14Nov 17, 2017Updated 8 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆45Feb 27, 2025Updated last year
- ☆11May 2, 2023Updated 3 years ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- Agent-native Seedance 2.0 short-film studio: cli for AI, canvas for human☆16Jun 14, 2026Updated 2 months ago
- Quantized Attention on GPU☆45Nov 22, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- 大模型部署实战:TensorRT-LLM, Triton Inference Server, vLLM☆27Feb 26, 2024Updated 2 years ago
- Desktop version of ChatGPT, support manually set cookie☆19Dec 9, 2022Updated 3 years ago
- This repository open-sources our GEC system submitted by THU KELab (sz) in the CCL2023-CLTC Track 1: Multidimensional Chinese Learner Tex…☆15Nov 25, 2023Updated 2 years ago
- An FPGA integration and acceleration of the popular FAISS framework for approximate similarity search☆25Jul 20, 2019Updated 7 years ago
- The driver for LMCache core to run in vLLM☆69Feb 4, 2025Updated last year
- WIP. Veloce is a low-code Ray-based parallelization library that makes machine learning computation novel, efficient, and heterogeneous.☆17Aug 4, 2022Updated 4 years ago
- semantic segmentation using keras☆15Apr 8, 2017Updated 9 years ago