WebAssembly (Wasm) Build and Bindings for llama.cpp
☆288Jul 23, 2024Updated last year
Alternatives and similar repositories for llama-cpp-wasm
Users that are interested in llama-cpp-wasm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inference☆1,029Dec 17, 2025Updated 3 months ago
- Cleanai (https://github.com/willmil11/cleanai) except I'm making it in c now. Fast and clean from the start this time :)☆17Mar 6, 2026Updated last month
- ☆19Feb 7, 2024Updated 2 years ago
- Semantic emoji finder. Python/dash UI. Uses sentence transformer embeddings and duckdb☆19Sep 15, 2025Updated 6 months ago
- Simple Tool Caller for llama.cpp☆11Aug 12, 2024Updated last year
- DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A collection of experiments related to LLM inference with llama.cpp/mlx☆40Updated this week
- Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation le…☆1,989Updated this week
- Tensor library for machine learning☆273Apr 23, 2023Updated 2 years ago
- Run Large-Language Models (LLMs) 🚀 directly in your browser!☆229Sep 8, 2024Updated last year
- Port of Microsoft's BioGPT in C/C++ using ggml☆86Feb 21, 2024Updated 2 years ago
- High-level, optionally asynchronous Rust bindings to llama.cpp☆245Jun 5, 2024Updated last year
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 7 months ago
- A Javascript library (with Typescript types) to parse metadata of GGML based GGUF files.☆52Jul 30, 2024Updated last year
- A live multiplayer trivia game where users can bid for the subject of the next question☆29Jan 9, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A super simple web interface to perform blind tests on LLM outputs.☆29Mar 9, 2024Updated 2 years ago
- ☆12Jun 27, 2024Updated last year
- Llama.cui is a small llama.cpp-based chat application for Node.js☆20Jul 10, 2025Updated 9 months ago
- High-performance In-browser LLM Inference Engine☆17,680Apr 3, 2026Updated last week
- Thin wrapper around GGML to make life easier☆45Nov 5, 2025Updated 5 months ago
- Web browser version of StarCoder.cpp☆46Jul 30, 2023Updated 2 years ago
- The rag pipeline for optimizing dynamic data editing.☆20Oct 30, 2025Updated 5 months ago
- Parallel wasm Barnes-Hut t-SNE implementation written in Rust.☆22Dec 27, 2025Updated 3 months ago
- Suno AI's Bark model in C/C++ for fast text-to-speech generation☆854Nov 16, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- llama explain chrome extension☆13Dec 28, 2023Updated 2 years ago
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated 11 months ago
- Minimalist stable-diffusion desktop application with only one executable file writen with golang ( No python ).☆18Apr 16, 2025Updated 11 months ago
- iterate quickly with llama.cpp hot reloading. use the llama.cpp bindings with bun.sh☆50Oct 30, 2023Updated 2 years ago
- Structured inference with Llama 2 in your browser☆52Nov 1, 2024Updated last year
- dart binding for llama.cpp☆288Jan 22, 2026Updated 2 months ago
- Demos for AI assistants using NLUX, Next.js, React, and Node.js☆17Jun 24, 2024Updated last year
- A cross-platform browser ML framework.☆753Apr 2, 2026Updated last week
- a pseudo-repo for discussion on Unix-like software in JS+Wasm ... and also about *browser* Python, Lua, Tcl.☆11Jan 25, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++☆5,682Apr 1, 2026Updated last week
- Python WASI build.☆10Jan 3, 2024Updated 2 years ago
- Inference Llama 2 in one file of pure C#☆24Sep 3, 2023Updated 2 years ago
- replacement of AdamW and Lion optimizer for LLMs☆13May 28, 2023Updated 2 years ago
- Downsampling array of intervals☆26Dec 11, 2019Updated 6 years ago
- Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.☆2,884Feb 10, 2026Updated 2 months ago
- Yet another `llama.cpp` Rust wrapper☆12Jun 19, 2024Updated last year