This is a repository accompanying the survey Edge AI Meets LLM (coming soon), containing a comprehensive list of papers, codebases, toolchains, and open-source frameworks. It is intended to serve as a handbook for researchers and developers interested in Edge/Mobile LLMs.
☆18Jun 5, 2025Updated last year
Alternatives and similar repositories for Awesome-Edge-LLMs
Users that are interested in Awesome-Edge-LLMs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Neural Network Quantization With Fractional Bit-widths☆11Feb 19, 2021Updated 5 years ago
- PyTorch Quantization Framework For OCP MX Datatypes.☆16May 30, 2025Updated last year
- Curated list of papers, frameworks, benchmarks, and applications for multimodal AI agents (LLMs, text-to-image, speech, world models, etc…☆39Updated this week
- ☆28Dec 2, 2024Updated last year
- An efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequences☆33Mar 7, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆19Dec 10, 2021Updated 4 years ago
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models☆31Apr 2, 2026Updated 5 months ago
- Dockerfiles for poetry/mlc-llm(rk3588)/...☆10Sep 13, 2023Updated 3 years ago
- Two-Stage ECG Signal Denoising Based Deep Convolutional Network☆13Nov 19, 2021Updated 4 years ago
- some simulation models of various processes (just for fun)☆10Feb 5, 2024Updated 2 years ago
- BEHRT: Transformer for Electronic Health Recrods☆12May 11, 2020Updated 6 years ago
- Run ops on Apple ANE in NPU register with pure python on M1 Asahi Linux. No Espresso, No CoreML, no metal, no .mlmodels file, no .hwx fil…☆19Jun 28, 2026Updated 2 months ago
- bitfusion verilog implementation☆13Feb 21, 2022Updated 4 years ago
- UniQL official repository (ICLR 2026)☆19Jan 27, 2026Updated 7 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Operating Systems Internals and Design principles 8th 读书笔记,资源整理☆19Jun 3, 2021Updated 5 years ago
- Awesome Blockchain Articles☆14Jul 10, 2022Updated 4 years ago
- ICLR 2021☆49Mar 18, 2021Updated 5 years ago
- Integrating Event-based Dynamic Vision Sensors with Sparse Hyperdimensional Computing☆13Jul 9, 2020Updated 6 years ago
- ☆20Apr 12, 2023Updated 3 years ago
- Deep Learning framework implemented from scratch in python using Numpy package.☆18Jan 28, 2020Updated 6 years ago
- Web chat front end for rk3588_npu_llm_server / RK3588 LLM chat interface☆16Jul 16, 2024Updated 2 years ago
- This is a 3-layer SNN for MNIST☆15May 11, 2022Updated 4 years ago
- Matrix multiplication on the NPU inside RK3588☆18Jun 27, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆10Mar 8, 2025Updated last year
- [TMLR] Official PyTorch implementation of paper "Efficient Quantization-aware Training with Adaptive Coreset Selection"☆39Aug 20, 2024Updated 2 years ago
- ☆28Aug 21, 2024Updated 2 years ago
- ☆124Nov 17, 2023Updated 2 years ago
- ☆14Apr 6, 2025Updated last year
- ☆15Oct 26, 2022Updated 3 years ago
- Qwen-Image's DiT inference with TensorRT-10☆21Oct 13, 2025Updated 11 months ago
- ☆14Feb 26, 2026Updated 6 months ago
- Reference book on Arm Helium (M-Profile Vector Extension) for Cortex-M processors covering SIMD, DSP, and ML (educational)☆26Jun 16, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆21Mar 2, 2026Updated 6 months ago
- ☆12Aug 18, 2023Updated 3 years ago
- Allows access via HTTP to LLM running on RK3588 NPU. Returns JSON response.☆30May 29, 2024Updated 2 years ago
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆26Sep 1, 2025Updated last year
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year
- ☆13Jun 29, 2024Updated 2 years ago
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆14Nov 27, 2024Updated last year