Community maintained hardware plugin for vLLM on AWS Neuron
☆47Aug 17, 2026Updated this week
Alternatives and similar repositories for vllm-neuron
Users that are interested in vllm-neuron are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆69Updated this week
- A Python-based tool, trained on the state-of-the-art Google Pegasus model, specializing in generating abstracts from given YouTube video …☆10Aug 6, 2023Updated 3 years ago
- ☆71Jun 22, 2026Updated last month
- A high-throughput and memory-efficient inference and serving engine for LLMs☆25Jun 9, 2026Updated 2 months ago
- ☆40Jul 28, 2026Updated 3 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆54May 19, 2025Updated last year
- SGLang kernel library for Intel XPU☆31Updated this week
- Spot Tagging Bot for Digital Assets☆18Jul 19, 2021Updated 5 years ago
- Collection of memory microbenchmarks to investigate NVIDIA GPUs Network on Chip architectures☆15Apr 14, 2026Updated 4 months ago
- ☆24Nov 18, 2025Updated 9 months ago
- A forked version of flux-fast that makes flux-fast even faster with cache-dit, 3.3x speedup on NVIDIA L20.☆24Jul 18, 2025Updated last year
- nnvm&tvm example of cross compilation and deployment in Nvidia Jetson TX2 platform☆11Apr 17, 2018Updated 8 years ago
- Demo for Qwen2.5-VL-3B-Instruct on Axera device.☆16Sep 3, 2025Updated 11 months ago
- A minimal hackable implementation of policy gradient methods (GRPO, PPO, REINFORCE)☆17Feb 20, 2026Updated 5 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- the symbol description of mobilenet v2☆11Sep 7, 2018Updated 7 years ago
- ☆16Aug 5, 2025Updated last year
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆30Dec 17, 2024Updated last year
- A safetensors extension to efficiently store sparse quantized tensors on disk☆313Updated this week
- ☆68Dec 29, 2025Updated 7 months ago
- ☆22May 5, 2025Updated last year
- [CVPR 2025] The official implementation of "CacheQuant: Comprehensively Accelerated Diffusion Models"☆48Nov 2, 2025Updated 9 months ago
- fake CUTLASS to get peformance☆26Apr 28, 2026Updated 3 months ago
- A DMA Controller for RISCV CPUs☆13Aug 10, 2015Updated 11 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- The (open-source part of) code to reproduce "BPPSA: Scaling Back-propagation by Parallel Scan Algorithm".☆13Jun 7, 2021Updated 5 years ago
- JAX support for tvm-ffi abi☆26May 14, 2026Updated 3 months ago
- The Amazon ECR Transfer Plugin for Data Transfer Hub(https://github.com/awslabs/data-transfer-hub). Transfer container images from Amazon…☆13Jan 29, 2025Updated last year
- Perf monitoring CLI tool for Apple Silicon☆10Jan 25, 2023Updated 3 years ago
- Auto Encoder on Tensorflow☆12Oct 18, 2017Updated 8 years ago
- A set of utilities to turn Dataclasses into useful configuration managers.☆11Mar 27, 2024Updated 2 years ago
- ☆23Jun 30, 2026Updated last month
- High-performance LLM operator library built on TileLang.☆169Updated this week
- ☆22Nov 6, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Utility functions/scripts for working with GPUs.☆10Jul 5, 2021Updated 5 years ago
- Minimal PyTorch implementation of TP, SP, FSDP and sharded-EMA☆32Nov 27, 2025Updated 8 months ago
- TPU inference for vLLM, with unified JAX and PyTorch support.☆410Updated this week
- ☆25Sep 4, 2025Updated 11 months ago
- 使用预训练语言模型ALBERT做中文NER☆12Jul 14, 2021Updated 5 years ago
- 👷 Build compute kernels☆213Apr 6, 2026Updated 4 months ago
- ☆121Aug 7, 2026Updated last week