Omni_Infer is a suite of inference accelerators designed for the Ascend NPU platform, offering native support and an expanding feature set.
☆127Jul 21, 2026Updated this week
Alternatives and similar repositories for omni-infer
Users that are interested in omni-infer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Community maintained hardware plugin for vLLM on Ascend☆2,450Updated this week
- SGLang kernel library for NPU☆166Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,139Updated this week
- LMCache on Ascend☆82Updated this week
- [ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"☆28Jun 1, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆5,941Updated this week
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆22May 5, 2026Updated 2 months ago
- A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Fou…☆1,479Updated this week
- AISBench Benchmark is a model evaluation tool built on OpenCompass, compatible with OpenCompass’s configuration system, dataset structure…☆169Updated this week
- ☆11Jan 26, 2022Updated 4 years ago
- Triton language and compiler for Ascend NPU☆112Updated this week
- ☆544Jul 14, 2026Updated last week
- Composable and Embeddable Communication Runtime for Distributed AI Services☆102Jun 5, 2026Updated last month
- KV cache store for distributed LLM inference