Manages vllm-nccl dependency
☆19Jun 3, 2024Updated 2 years ago
Alternatives and similar repositories for vllm-nccl
Users that are interested in vllm-nccl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PSTensor provides a way to hack the memory management of tensors in TensorFlow and PyTorch by defining your own C++ Tensor Class.☆10Feb 10, 2022Updated 4 years ago
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Sep 11, 2026Updated 2 weeks ago
- An auxiliary project analysis of the characteristics of KV in DiT Attention.☆34Nov 29, 2024Updated last year
- ☆20May 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Elixir: Train a Large Language Model on a Small GPU Cluster☆15Jun 8, 2023Updated 3 years ago
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- This project is based on the [LTX-Video](https://github.com/Lightricks/LTX-Video) algorithm of the diffusers and optimized and accelerate…☆16Dec 31, 2024Updated last year
- A Triton-only attention backend for vLLM☆28Jul 14, 2026Updated 2 months ago
- ☆21Mar 22, 2021Updated 5 years ago
- A Feishu/Lark AI agent bot☆16Feb 27, 2026Updated 7 months ago
- ☆13Jul 24, 2024Updated 2 years ago
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- SmartNIC☆14Dec 13, 2018Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆15Dec 1, 2023Updated 2 years ago
- Debug DeepSpeed-Chat step by step in IDE (在IDE里一步一步调试DeepSpeed-Chat)☆10Apr 17, 2023Updated 3 years ago
- A Translation Task using TurboTransformers☆10Dec 17, 2020Updated 5 years ago
- ☆20Sep 28, 2024Updated 2 years ago
- ☆28Jul 29, 2025Updated last year
- This is an official GitHub repository for the paper, "Towards timeout-less transport in commodity datacenter networks.".☆17Oct 12, 2021Updated 4 years ago
- 3位代码类目表;6位扩展代码表;疾病分类与代码(修订版);章节名称及代码☆11Aug 20, 2018Updated 8 years ago
- Drop-in library for tracking the memory allocations of CUDA applications☆14Nov 17, 2017Updated 8 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆46Feb 27, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆11May 2, 2023Updated 3 years ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- creditmodel, 模型,评分卡,scorecard, vintage, automatic modeling☆11Aug 10, 2024Updated 2 years ago
- Quantized Attention on GPU☆45Nov 22, 2024Updated last year
- An FPGA integration and acceleration of the popular FAISS framework for approximate similarity search☆25Jul 20, 2019Updated 7 years ago
- WIP. Veloce is a low-code Ray-based parallelization library that makes machine learning computation novel, efficient, and heterogeneous.☆17Aug 4, 2022Updated 4 years ago
- semantic segmentation using keras☆15Apr 8, 2017Updated 9 years ago
- ☆28Mar 2, 2023Updated 3 years ago
- Exploration of semantic chunking and chunk classification☆19Sep 16, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity☆247Sep 24, 2023Updated 3 years ago
- ☆13Oct 26, 2020Updated 5 years ago
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attention☆57Aug 6, 2025Updated last year
- [MLSys 2026] AccelOpt: Self-improving Agents for AI Accelerator Kernel Optimization☆75Jul 28, 2026Updated last month
- Manually implemented quantization-aware training☆22Oct 12, 2022Updated 3 years ago
- Network Traffic Transformer to learn network dynamics from packet traces. Learn fundamental dynamics with pre-training and fine-tune to m…☆24Jan 17, 2024Updated 2 years ago
- ☆23May 6, 2022Updated 4 years ago