WebAssembly (Wasm) Build and Bindings for llama.cpp
☆293Jul 23, 2024Updated 2 years ago
Alternatives and similar repositories for llama-cpp-wasm
Users that are interested in llama-cpp-wasm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inference☆1,162Jun 17, 2026Updated last month
- A collection of experiments related to LLM inference with llama.cpp/mlx☆40Updated this week
- Simple Tool Caller for llama.cpp☆11Aug 12, 2024Updated last year
- Run Large-Language Models (LLMs) 🚀 directly in your browser!☆233Sep 8, 2024Updated last year
- Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation le…☆2,152Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Tensor library for machine learning☆273Apr 23, 2023Updated 3 years ago
- Inference of Mamba, Mamba2 and Mamba3 models in pure C☆203Mar 18, 2026Updated 4 months ago
- Cleanai (https://github.com/willmil11/cleanai) except I'm making it in c now. Fast and clean from the start this time :)☆15Jul 17, 2026Updated 3 weeks ago
- Port of Microsoft's BioGPT in C/C++ using ggml☆87Feb 21, 2024Updated 2 years ago
- GGML implementation of BERT model with Python bindings and quantization.☆57Feb 19, 2024Updated 2 years ago
- High-level, optionally asynchronous Rust bindings to llama.cpp☆251Jun 5, 2024Updated 2 years ago
- A live multiplayer trivia game where users can bid for the subject of the next question☆29Jan 9, 2026Updated 7 months ago
- High-performance In-browser LLM Inference Engine☆18,541Updated this week
- Llama.cui is a small llama.cpp-based chat application for Node.js☆19Jul 10, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- LLM plugin for interacting with llama-server models☆32May 28, 2025Updated last year
- Thin wrapper around GGML to make life easier☆49Jul 26, 2026Updated 2 weeks ago
- A Javascript library (with Typescript types) to parse metadata of GGML based GGUF files☆52Jul 30, 2024Updated 2 years ago
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- Suno AI's Bark model in C/C++ for fast text-to-speech generation☆866Nov 16, 2024Updated last year
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 11 months ago
- iterate quickly with llama.cpp hot reloading. use the llama.cpp bindings with bun.sh☆51Oct 30, 2023Updated 2 years ago
- Demos for AI assistants using NLUX, Next.js, React, and Node.js☆17Jun 24, 2024Updated 2 years ago
- A super simple web interface to perform blind tests on LLM outputs.☆30Mar 9, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A JavaScript to WebAssembly transpiler☆40Oct 28, 2023Updated 2 years ago
- Structured inference with Llama 2 in your browser☆52Nov 1, 2024Updated last year
- A cross-platform browser ML framework.☆768May 26, 2026Updated 2 months ago
- Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)☆915Updated this week
- replacement of AdamW and Lion optimizer for LLMs☆13May 28, 2023Updated 3 years ago
- State-of-the-art Machine Learning for the web. Run 🤗 Transformers directly in your browser, with no need for a server!☆16,241Jul 31, 2026Updated last week
- Implementation of YOLO (You Only Look Once) computer Vision algorithm in a React UI, for the subject Intelligent Systems (ULL)☆11Jan 27, 2019Updated 7 years ago
- Some plugins to help you unlock new functionality in Selenium IDE☆12Dec 7, 2022Updated 3 years ago
- ☆48Mar 9, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LLama.cpp rust bindings☆424Jun 27, 2024Updated 2 years ago
- Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.☆3,032Jul 5, 2026Updated last month
- Syntexmex plugin for blender☆16Mar 28, 2020Updated 6 years ago
- Easily convert HuggingFace models to GGUF-format for llama.cpp☆22Jul 27, 2024Updated 2 years ago
- Zero-shot forecasting, tabular classification, and regression via MCP — exposes Google TimesFM 2.5 and TabFM v1.0.0 to AI assistants. Jus…☆27Jul 12, 2026Updated 3 weeks ago
- Experiments on speculative sampling with Llama models☆129Jun 8, 2023Updated 3 years ago
- Vercel and web-llm template to run wasm models directly in the browser.☆180Apr 17, 2026Updated 3 months ago