Community maintained hardware plugin for vLLM on AWS Neuron
☆50Aug 17, 2026Updated last month
Alternatives and similar repositories for vllm-neuron
Users that are interested in vllm-neuron are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆72Sep 23, 2026Updated last week
- A Python-based tool, trained on the state-of-the-art Google Pegasus model, specializing in generating abstracts from given YouTube video …☆10Aug 6, 2023Updated 3 years ago
- ☆72Sep 22, 2026Updated last week
- Show the time in Roman Numerals☆12Jan 23, 2020Updated 6 years ago
- ☆40Jul 28, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆54May 19, 2025Updated last year
- SGLang kernel library for Intel XPU☆36Updated this week
- Spot Tagging Bot for Digital Assets☆18Jul 19, 2021Updated 5 years ago
- Powering AWS purpose-built machine learning chips. Blazing fast and cost effective, natively integrated into PyTorch and TensorFlow and i…☆637Updated this week
- ☆24Nov 18, 2025Updated 10 months ago
- A forked version of flux-fast that makes flux-fast even faster with cache-dit, 3.3x speedup on NVIDIA L20.☆24Jul 18, 2025Updated last year
- Intel Gaudi's Megatron DeepSpeed Large Language Models for training☆17Dec 19, 2024Updated last year
- nnvm&tvm example of cross compilation and deployment in Nvidia Jetson TX2 platform☆11Apr 17, 2018Updated 8 years ago
- Demo for Qwen2.5-VL-3B-Instruct on Axera device.☆16Sep 3, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository contains code used for our Multi Sentence Inference NAACL'22 paper.☆12Mar 6, 2023Updated 3 years ago
- A minimal hackable implementation of policy gradient methods (GRPO, PPO, REINFORCE)☆18Feb 20, 2026Updated 7 months ago
- Community maintained hardware plugin for vLLM on Spyre☆53Updated this week
- Evaluate Transformers from the Hub 🔥☆14Sep 24, 2026Updated last week
- ☆16Aug 5, 2025Updated last year
- A safetensors extension to efficiently store sparse quantized tensors on disk☆329Updated this week
- [CVPR 2025] The official implementation of "CacheQuant: Comprehensively Accelerated Diffusion Models"☆51Nov 2, 2025Updated 11 months ago
- fake CUTLASS to get peformance☆25Apr 28, 2026Updated 5 months ago
- A DMA Controller for RISCV CPUs☆13Aug 10, 2015Updated 11 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- NKIPy: Rapid Prototyping on Trainium☆30Sep 14, 2026Updated 2 weeks ago
- JAX support for tvm-ffi abi☆28May 14, 2026Updated 4 months ago
- The Amazon ECR Transfer Plugin for Data Transfer Hub(https://github.com/awslabs/data-transfer-hub). Transfer container images from Amazon…☆13Jan 29, 2025Updated last year
- ☆13Sep 10, 2026Updated 3 weeks ago
- A set of utilities to turn Dataclasses into useful configuration managers.☆11Mar 27, 2024Updated 2 years ago
- ☆23Jun 30, 2026Updated 3 months ago
- High-performance LLM operator library built on TileLang.☆191Updated this week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆58Updated this week
- Utility functions/scripts for working with GPUs.☆10Jul 5, 2021Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Minimal PyTorch implementation of TP, SP, FSDP and sharded-EMA☆32Nov 27, 2025Updated 10 months ago
- TPU inference for vLLM, with unified JAX and PyTorch support.☆450Updated this week
- ☆26Sep 4, 2025Updated last year
- 使用预训练语言模型ALBERT做中文NER☆12Jul 14, 2021Updated 5 years ago
- 👷 Build compute kernels☆213Apr 6, 2026Updated 5 months ago
- Memory experiments with LLMs☆11Mar 31, 2023Updated 3 years ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 5 months ago