A collection of experiments related to LLM inference with llama.cpp/mlx
☆40Jun 18, 2026Updated last month
Alternatives and similar repositories for llama-sandbox
Users that are interested in llama-sandbox are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ffmpeg+cuvid+tensorrt+multicamera☆12Dec 31, 2024Updated last year
- Testing LLM reasoning abilities with family relationship quizzes.☆62Jan 28, 2025Updated last year
- ☆31Dec 23, 2024Updated last year
- ☆18Dec 7, 2023Updated 2 years ago
- Build tools for Enyo 2.6+☆15Nov 26, 2018Updated 7 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- minimal C implementation of speculative decoding based on llama2.c☆30Jul 15, 2024Updated 2 years ago
- LLM plugin for interacting with llama-server models☆32May 28, 2025Updated last year
- Inference of Mamba, Mamba2 and Mamba3 models in pure C☆203Mar 18, 2026Updated 4 months ago
- Documentation for Enyo and its libraries☆16May 9, 2020Updated 6 years ago
- Recording models☆13Sep 19, 2023Updated 2 years ago
- 本仓库在OpenVINO推理框架下部署Nanodet检测算法,并重写预处理和后处理部分,具有超高性能!让你在Intel CPU平台上的检测速度起飞! 并基于NNCF和PPQ工具将模型量化(PTQ)至int8精度,推理速度更快!☆16Jun 14, 2023Updated 3 years ago
- A C++17 single-file header-only wrapper for llama.cpp☆30Updated this week
- a bower registry written in python and django☆16Nov 16, 2014Updated 11 years ago
- Repository for ICLR'23 Long-tailed Learning Requires Feature Learning☆10Feb 22, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆23Apr 10, 2024Updated 2 years ago
- A live multiplayer trivia game where users can bid for the subject of the next question☆29Jan 9, 2026Updated 6 months ago
- ☆20Dec 29, 2023Updated 2 years ago
- Light weight, ultra portable, responsive javascript library, responsive all the things.☆20Jan 28, 2015Updated 11 years ago
- monodepth running in Android by ncnn☆23Oct 12, 2021Updated 4 years ago
- Downsampling array of intervals☆26Dec 11, 2019Updated 6 years ago
- Proof-of-concept of global switching between numpy/jax/pytorch in a library.☆17Jun 18, 2024Updated 2 years ago
- 📊 an html tap reporter☆20Sep 7, 2023Updated 2 years ago
- UI infrastructure for Enyo applications☆20Mar 26, 2018Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for CIKM 2021 best short paper nomination "Modeling Sequences as Distributions with Uncertainty for Sequential Recommendation" https…☆16Jun 11, 2021Updated 5 years ago
- LLM-powered lossless compression tool☆317Jun 16, 2026Updated last month
- Semantic emoji finder. Python/dash UI. Uses sentence transformer embeddings and duckdb☆20Sep 15, 2025Updated 10 months ago
- ☆18Jun 16, 2025Updated last year
- Simple ranking metrics for PyTorch on CPU or GPU☆15Nov 20, 2020Updated 5 years ago
- ☆34Jul 23, 2024Updated 2 years ago
- Llama3 Streaming Chat Sample☆22Apr 24, 2024Updated 2 years ago
- TypeScript generator for llama.cpp Grammar directly from TypeScript interfaces☆146Jul 9, 2024Updated 2 years ago
- Decrypt multicast Verimatrix streams☆14Apr 21, 2022Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Pybind11 bindings for Whisper.cpp☆63Updated this week
- ☆16Jul 11, 2025Updated last year
- Really, really lightweight javascript event emitting☆22Jan 24, 2013Updated 13 years ago
- WebAssembly (Wasm) Build and Bindings for llama.cpp☆293Jul 23, 2024Updated 2 years ago
- [ICML 2022] Official implementation of "Score-Guided Intermediate Layer Optimization: Fast Langevin Mixing for Inverse Problems".☆12Jul 19, 2022Updated 4 years ago
- A tool convert TensorRT engine/plan to a fake onnx☆41Nov 22, 2022Updated 3 years ago
- ☆30Nov 16, 2024Updated last year