vLLM plugin for RBLN NPU
☆59Sep 19, 2026Updated this week
Alternatives and similar repositories for vllm-rbln
Users that are interested in vllm-rbln are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ⚡ A seamless integration of HuggingFace Transformers & Diffusers with RBLN SDK for efficient inference on RBLN NPUs.☆19Updated this week
- Ditto is an open-source framework that enables direct conversion of HuggingFace PreTrainedModels into TensorRT-LLM engines.☆56Jul 16, 2025Updated last year
- QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference☆123Mar 6, 2024Updated 2 years ago
- RBLN Model Zoo — Compile once. Deploy anywhere.☆40Aug 28, 2026Updated 3 weeks ago
- Provides the examples to write and build Habana custom kernels using the HabanaTools☆26Apr 15, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Layered prefill changes the scheduling axis from tokens to layers and removes redundant MoE weight reloads while keeping decode stall fre…☆20Mar 9, 2026Updated 6 months ago
- ☆55Nov 22, 2022Updated 3 years ago
- Source code for Trinity(ASPLOS 2026)☆26Apr 24, 2026Updated 4 months ago
- Dashboard for InferenceX™, Open Source Continuous Inference | InferenceX 仪表板☆45Updated this week
- ☆11Aug 23, 2023Updated 3 years ago
- Community maintained hardware plugin for vLLM on Spyre☆53Sep 9, 2026Updated last week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆57Updated this week
- Experimental projects related to TensorRT☆127Jul 2, 2026Updated 2 months ago
- GenCoG: A DSL-Based Approach to Generating Computation Graphs for TVM Testing (ISSTA‘23)☆17Jul 19, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Sep 11, 2026Updated last week
- minimal C implementation of speculative decoding based on llama2.c☆30Jul 15, 2024Updated 2 years ago
- ☆107Sep 9, 2024Updated 2 years ago
- ☆15Jan 7, 2023Updated 3 years ago
- A simple GPU reservation tool for single host shared development systems☆30Aug 24, 2026Updated 3 weeks ago
- ☆171Feb 15, 2025Updated last year
- A tool for decomposing DRAM address mapping into component-level functions☆17Jun 12, 2025Updated last year
- [CVPR 2023] "TrojViT: Trojan Insertion in Vision Transformers" by Mengxin Zheng, Qian Lou, Lei Jiang☆15Jan 5, 2024Updated 2 years ago
- Panopticon is a complete in-DRAM RowHammer mitigation. This code simulates an implementation of Panopticon in DDR5.☆14Jun 2, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- extensible collectives library in triton☆97Mar 31, 2025Updated last year
- ☆10Apr 24, 2023Updated 3 years ago
- ☆27Jan 8, 2024Updated 2 years ago
- ☆17Apr 9, 2025Updated last year
- ☆15Mar 30, 2024Updated 2 years ago
- ☆39Dec 14, 2025Updated 9 months ago
- ☆20Sep 6, 2026Updated 2 weeks ago
- TPU inference for vLLM, with unified JAX and PyTorch support.☆435Updated this week
- Performant kernels for symmetric tensors☆17Aug 22, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Sep 10, 2026Updated last week
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- ☆69Dec 29, 2025Updated 8 months ago
- SParse AcceleRation on Tensor Architecture☆18Apr 15, 2026Updated 5 months ago
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.☆14Jan 8, 2026Updated 8 months ago
- Bluespec environment for working with the ulx3s board and its lattice ecp5 fpga☆15Mar 9, 2025Updated last year
- Pytorch distributed backend extension with compression support☆17Mar 24, 2025Updated last year