Community maintained hardware plugin for vLLM on AWS Neuron
☆51Aug 17, 2026Updated 3 weeks ago
Alternatives and similar repositories for vllm-neuron
Users that are interested in vllm-neuron are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆70Updated this week
- ☆72Aug 26, 2026Updated last week
- Show the time in Roman Numerals☆12Jan 23, 2020Updated 6 years ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆24Jun 9, 2026Updated 2 months ago
- ☆40Jul 28, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆54May 19, 2025Updated last year
- SGLang kernel library for Intel XPU☆32Updated this week
- Spot Tagging Bot for Digital Assets☆18Jul 19, 2021Updated 5 years ago
- Powering AWS purpose-built machine learning chips. Blazing fast and cost effective, natively integrated into PyTorch and TensorFlow and i…☆631Aug 26, 2026Updated last week
- ☆24Nov 18, 2025Updated 9 months ago
- A forked version of flux-fast that makes flux-fast even faster with cache-dit, 3.3x speedup on NVIDIA L20.☆24Jul 18, 2025Updated last year
- Community maintained hardware plugin for vLLM on Spyre☆53Aug 31, 2026Updated last week
- the symbol description of mobilenet v2☆11Sep 7, 2018Updated 8 years ago
- Evaluate Transformers from the Hub 🔥☆14May 26, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆16Aug 5, 2025Updated last year
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆30Dec 17, 2024Updated last year
- A safetensors extension to efficiently store sparse quantized tensors on disk☆322Updated this week
- ☆69Dec 29, 2025Updated 8 months ago
- NKIPy: Rapid Prototyping on Trainium☆30Updated this week
- [CVPR 2025] The official implementation of "CacheQuant: Comprehensively Accelerated Diffusion Models"☆51Nov 2, 2025Updated 10 months ago
- fake CUTLASS to get peformance☆26Apr 28, 2026Updated 4 months ago
- The (open-source part of) code to reproduce "BPPSA: Scaling Back-propagation by Parallel Scan Algorithm".☆13Jun 7, 2021Updated 5 years ago
- JAX support for tvm-ffi abi☆27May 14, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Aug 12, 2022Updated 4 years ago
- Perf monitoring CLI tool for Apple Silicon☆10Jan 25, 2023Updated 3 years ago
- ☆13Aug 18, 2026Updated 3 weeks ago
- A set of utilities to turn Dataclasses into useful configuration managers.☆11Mar 27, 2024Updated 2 years ago
- ☆23Jun 30, 2026Updated 2 months ago
- High-performance LLM operator library built on TileLang.☆178Updated this week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆53Updated this week
- ☆22Nov 6, 2025Updated 10 months ago
- ☆11Aug 27, 2019Updated 7 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Utility functions/scripts for working with GPUs.☆10Jul 5, 2021Updated 5 years ago
- Minimal PyTorch implementation of TP, SP, FSDP and sharded-EMA☆32Nov 27, 2025Updated 9 months ago
- Coala is a python package for Contextual Answer Sentence Selection.☆15Jun 12, 2023Updated 3 years ago
- TPU inference for vLLM, with unified JAX and PyTorch support.☆425Updated this week
- Deep Learning inference with AWS Lambda and Amazon EFS☆14Aug 24, 2020Updated 6 years ago
- ☆25Sep 4, 2025Updated last year
- An Example of MXNet Models Comilation and Deployment with NNVM in C++☆16Apr 25, 2018Updated 8 years ago