☆29Aug 3, 2026Updated last week
Alternatives and similar repositories for llama-cpp-qnn-builder
Users that are interested in llama-cpp-qnn-builder are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM inference in C/C++☆53Aug 3, 2026Updated last week
- the original reference implementation of a specified llama.cpp backend for Qualcomm Hexagon NPU on Android phone, history of ggml-hexagon…☆50Updated this week
- LLM inference in C/C++☆22Oct 22, 2025Updated 9 months ago
- snpe tutorial☆10Dec 25, 2023Updated 2 years ago
- Inference of YOLOv7 model applied on Qualcomm SNPE for pedestrian detection with embedded system.☆14Sep 23, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A CUDA kernel for NHWC GroupNorm for PyTorch☆25Nov 15, 2024Updated last year
- Hopenet: deep head pose estimator on ncnn☆10Jun 18, 2020Updated 6 years ago
- ☆11Feb 5, 2026Updated 6 months ago
- Ultra fast head pose estimation on a bare Raspberry Pi 4 at 20 FPS☆10Dec 21, 2021Updated 4 years ago
- High-speed and easy-use LLM serving framework for local deployment☆165Aug 7, 2025Updated last year
- High-performance system monitor for Rockchip SoCs (RK3588, RK3399) with real-time CPU, GPU, NPU, RGA, memory, and process monitoring. W…☆28Nov 25, 2025Updated 8 months ago
- onnxruntime-qnn is the Qualcomm AI Runtime (QAIRT) execution provider for onnxruntime. It provides onnxruntime hardware acceleration and …☆43Updated this week
- Optimized pose detector inference for edge devices☆15Feb 23, 2023Updated 3 years ago
- FastRPC is Qualcomm's userspace library that facilitates efficient remote procedure calls between the CPU and DSP for high-performance co…☆107Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- hexagon tutorial☆58Mar 29, 2026Updated 4 months ago
- Self-implemented NN operators for Qualcomm's Hexagon NPU☆77Sep 30, 2025Updated 10 months ago
- A Android Library for YOLOv5/YOLOv7/YOLOv8 Detection and Pose Inference Based on NCNN☆59Aug 15, 2024Updated last year
- 智能家教微信小程序☆11Sep 15, 2018Updated 7 years ago
- A rust version of the Caffe library.☆19Jun 16, 2021Updated 5 years ago
- End to End Speech to Speech with Emotion System☆15Feb 6, 2025Updated last year
- 使用onnxruntime部署Gaze-LLE凝视目标估计,包含C++和Python两个版本的程序☆17Jan 21, 2025Updated last year
- Python utility to convert PyTorch model weights from '.bin' to '.safetensors' format.☆18Sep 19, 2025Updated 10 months ago
- this is just a copy of https://code.msdn.microsoft.com/windowsdesktop/DirectCompute-Graphics-425de5a8☆11Nov 1, 2017Updated 8 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Apr 28, 2023Updated 3 years ago
- 官方transformers源码解析。AI大模型时代,pytorch、transformer是新操作系统,其他都是运行在其上面的软件。☆16Sep 25, 2023Updated 2 years ago
- CUDA SGEMM optimization note☆15Oct 31, 2023Updated 2 years ago
- ☆11May 19, 2025Updated last year
- PyCon mini 東海 2024 のトーク「Google Colaboratoryで試すVLM 」で紹介したサンプル集☆12Nov 15, 2024Updated last year
- NetHCF: Enabling Line-rate and Adaptive Spoofed IP Traffic Filtering☆13Mar 17, 2022Updated 4 years ago
- Acclaim: Adaptive Memory Reclaim to Improve User Experience in Android Systems [ATC '20]☆16Aug 1, 2020Updated 6 years ago
- Build environment for Raspberry Pi Pico (RP2040) C/C++ SDK☆11Jan 25, 2021Updated 5 years ago
- [ACL2025 Oral🔥]Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling☆29Nov 11, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for paper "ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection" (MobiSys'23)☆14Nov 1, 2023Updated 2 years ago
- ☆13Mar 18, 2024Updated 2 years ago
- Artifacts of VLDB'22 paper "COMET: A Novel Memory-Efficient Deep Learning TrainingFramework by Using Error-Bounded Lossy Compression"☆10Aug 2, 2022Updated 4 years ago
- RAG-QA is a free, containerised question-answer framework that allows you to ask questions to your documents in an intuitive way☆21Jan 25, 2024Updated 2 years ago
- This is the Pytorch implementation of paper--Training deep neural-networks using a noise adaptation layer.☆10Apr 18, 2021Updated 5 years ago
- Flash Attention in raw Cuda C beating PyTorch☆39May 14, 2024Updated 2 years ago
- Bjontegaard metric computation in the python language☆17Oct 9, 2020Updated 5 years ago