High-speed and easy-use LLM serving framework for local deployment
☆165Aug 7, 2025Updated last year
Alternatives and similar repositories for PowerServe
Users that are interested in PowerServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bamboo-7B Large Language Model☆94Mar 28, 2024Updated 2 years ago
- ☆93Dec 16, 2025Updated 7 months ago
- ☆51Jul 30, 2025Updated last year
- Inference RWKV v5, v6 and v7 with Qualcomm AI Engine Direct SDK☆95Jul 27, 2026Updated 2 weeks ago
- the original reference implementation of a specified llama.cpp backend for Qualcomm Hexagon NPU on Android phone, history of ggml-hexagon…☆51Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Fast Multimodal LLM on Mobile Devices☆1,587Updated this week
- YOLOv5在高通AI Engine Direct环境下进行QNN量化,CPU推理的项目☆17Sep 10, 2024Updated last year
- ☆29Aug 3, 2026Updated last week
- hexagon tutorial☆61Mar 29, 2026Updated 4 months ago
- onnxruntime-qnn is the Qualcomm AI Runtime (QAIRT) execution provider for onnxruntime. It provides onnxruntime hardware acceleration and …☆43Updated this week
- [EMNLP Findings 2024] MobileQuant: Mobile-friendly Quantization for On-device Language Models☆69Sep 22, 2024Updated last year
- The Qualcomm® AI Hub apps are a collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) a…☆444Updated this week
- Source code of the paper "V-Droid: Advancing Mobile GUI Agent Through Generative Verifiers"☆35Feb 2, 2026Updated 6 months ago
- ☆11May 19, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- QAI AppBuilder is designed to help developers easily execute models on WoS and Linux platforms. It encapsulates the Qualcomm® AI Runtime …☆194Updated this week
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆23Jun 23, 2026Updated last month
- LLM inference in C/C++☆53Aug 3, 2026Updated last week
- High-speed Large Language Model Serving for Local Deployment☆9,715May 11, 2026Updated 3 months ago
- ☆10Jun 25, 2026Updated last month
- Mic-controlled mouse clicks☆17Oct 6, 2025Updated 10 months ago
- ☆181Jun 21, 2026Updated last month
- Let's use Qualcomm NPU in Android☆21Feb 18, 2025Updated last year
- 本项目是一个通过文字生成图片的项目,基于开源模型Stable Diffusion V1.5生成可以在手机的CPU和NPU上运行的模型,包括其配套的模型运行框架。☆247Mar 29, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 4 months ago
- The open-source materials for paper "Sparsing Law: Towards Large Language Models with Greater Activation Sparsity".☆32Nov 12, 2024Updated last year
- Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) an…☆1,185Updated this week
- ☆202Updated this week
- ☆23Jun 14, 2023Updated 3 years ago
- Code for paper "ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection" (MobiSys'23)☆14Nov 1, 2023Updated 2 years ago
- C++ implementations for various tokenizers (sentencepiece, tiktoken etc).☆50Aug 5, 2026Updated last week
- Pure C wrapper library to use llama.cpp with Linux and Windows as simple as possible.☆15Jul 28, 2026Updated 2 weeks ago
- An fully autonomous agent that accesses the browser and performs tasks.☆18Apr 25, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- Low-bit LLM inference on CPU/NPU with lookup table☆983Jun 5, 2025Updated last year
- [arXiv] On-device Sora: Enabling Diffusion-Based Text-to-Video Generation for Mobile Devices☆139Nov 27, 2025Updated 8 months ago
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆24Nov 26, 2025Updated 8 months ago
- KV cache compression for high-throughput LLM inference☆159Feb 5, 2025Updated last year
- This project is intended to build and deploy an SNPE model on Qualcomm Devices, which are having unsupported layers which are not part of…☆10Oct 4, 2021Updated 4 years ago
- ☆36Apr 2, 2026Updated 4 months ago