☆31Aug 3, 2026Updated 3 weeks ago
Alternatives and similar repositories for llama-cpp-qnn-builder
Users that are interested in llama-cpp-qnn-builder are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM inference in C/C++☆54Aug 24, 2026Updated last week
- the original FastRPC-based implementation of a specified llama.cpp backend for Qualcomm Hexagon NPU, history of ggml-hexagon: https://git…☆54Updated this week
- LLM inference in C/C++☆22Oct 22, 2025Updated 10 months ago
- snpe tutorial☆10Dec 25, 2023Updated 2 years ago
- Inference of YOLOv7 model applied on Qualcomm SNPE for pedestrian detection with embedded system.☆14Sep 23, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A CUDA kernel for NHWC GroupNorm for PyTorch☆26Nov 15, 2024Updated last year
- Hopenet: deep head pose estimator on ncnn☆10Jun 18, 2020Updated 6 years ago
- ☆11Feb 5, 2026Updated 6 months ago
- Ultra fast head pose estimation on a bare Raspberry Pi 4 at 20 FPS☆10Dec 21, 2021Updated 4 years ago
- High-speed and easy-use LLM serving framework for local deployment☆166Aug 7, 2025Updated last year
- High-performance system monitor for Rockchip SoCs (RK3588, RK3399) with real-time CPU, GPU, NPU, RGA, memory, and process monitoring. W…☆29Nov 25, 2025Updated 9 months ago
- onnxruntime-qnn is the Qualcomm AI Runtime (QAIRT) execution provider for onnxruntime. It provides onnxruntime hardware acceleration and …☆45Updated this week
- FastRPC is Qualcomm's userspace library that facilitates efficient remote procedure calls between the CPU and DSP for high-performance co…☆111Updated this week
- QAI AppBuilder is designed to help developers easily execute models on WoS and Linux platforms. It encapsulates the Qualcomm® AI Runtime …☆201Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- hexagon tutorial☆66Mar 29, 2026Updated 5 months ago
- ☆13Oct 17, 2024Updated last year
- Shadowsocks/ShadowsocksR 账号在线监控☆12Nov 25, 2018Updated 7 years ago
- A Android Library for YOLOv5/YOLOv7/YOLOv8 Detection and Pose Inference Based on NCNN☆60Aug 15, 2024Updated 2 years ago
- deepstream + cuda,yolo26,yolo-master,yolo11,yolov8,sam,transformer, etc.☆27Feb 7, 2026Updated 6 months ago
- 智能家教微信小程序☆11Sep 15, 2018Updated 7 years ago
- End to End Speech to Speech with Emotion System☆15Feb 6, 2025Updated last year
- ☆17Apr 9, 2025Updated last year
- 使用onnxruntime部署Gaze-LLE凝视目标估计,包含C++和Python两个版本的程序☆17Jan 21, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆15Apr 28, 2023Updated 3 years ago
- CUDA SGEMM optimization note☆15Oct 31, 2023Updated 2 years ago
- An example app of DNNLibrary :)☆13Jul 26, 2019Updated 7 years ago
- ☆12May 19, 2025Updated last year
- PyCon mini 東海 2024 のトーク「Google Colaboratoryで試すVLM」で紹介したサンプル集☆12Nov 15, 2024Updated last year
- NetHCF: Enabling Line-rate and Adaptive Spoofed IP Traffic Filtering☆13Mar 17, 2022Updated 4 years ago
- A graph coloring register allocator for LLVM.☆11Jan 23, 2017Updated 9 years ago
- Acclaim: Adaptive Memory Reclaim to Improve User Experience in Android Systems [ATC '20]☆16Aug 1, 2020Updated 6 years ago
- 🎵 Control YouTube players with browser by Alfred☆12Sep 30, 2020Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ACL2025 Oral🔥]Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling☆30Nov 11, 2025Updated 9 months ago
- 本项目是一个通过文字生成图片的项目,基于开源模型Stable Diffusion V1.5生成可以在手机的CPU和NPU上运行的模型,包括其配套的模型运行框架。☆248Mar 29, 2024Updated 2 years ago
- ☆13Mar 18, 2024Updated 2 years ago
- Artifacts of VLDB'22 paper "COMET: A Novel Memory-Efficient Deep Learning TrainingFramework by Using Error-Bounded Lossy Compression"☆10Aug 2, 2022Updated 4 years ago
- RAG-QA is a free, containerised question-answer framework that allows you to ask questions to your documents in an intuitive way☆21Jan 25, 2024Updated 2 years ago
- Flash Attention in raw Cuda C beating PyTorch☆39May 14, 2024Updated 2 years ago
- This project is intended to build and deploy an SNPE model on Qualcomm Devices, which are having unsupported layers which are not part of…☆10Oct 4, 2021Updated 4 years ago