Rust+OpenCL+AVX2 implementation of LLaMA inference code
☆554Feb 12, 2024Updated 2 years ago
Alternatives and similar repositories for rllama
Users that are interested in rllama are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [Unmaintained, see README] An ecosystem of Rust libraries for working with large language models☆6,147Jun 24, 2024Updated last year
- A relatively basic implementation of RWKV in Rust written by someone with very little math and ML knowledge. Supports 32, 8 and 4 bit eva…☆94Sep 2, 2023Updated 2 years ago
- Deep learning in Rust, with shape checked tensors and neural networks☆1,904Jul 23, 2024Updated last year
- A fast llama2 decoder in pure Rust.☆1,062Nov 30, 2023Updated 2 years ago
- Bleeding edge low level Rust binding for GGML☆17Jun 26, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- `llm-chain` is a powerful rust crate for building chains in large language models allowing you to summarise text and complete complex tas…☆1,600Oct 31, 2024Updated last year
- An implementation of the diffusers api in Rust☆591Apr 4, 2024Updated 2 years ago
- Inference Llama 2 in one file of pure Rust 🦀☆235Sep 11, 2023Updated 2 years ago
- LLaMa 7b with CUDA acceleration implemented in rust. Minimal GPU memory needed!☆111Jul 27, 2023Updated 2 years ago
- Rust bindings for the C++ api of PyTorch.☆5,392May 17, 2026Updated last week
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.☆15,223Updated this week
- ☆58Apr 6, 2023Updated 3 years ago
- LLama.cpp rust bindings☆423Jun 27, 2024Updated last year
- Llama2 LLM ported to Rust burn☆280Apr 16, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Minimalist ML framework for Rust☆20,324Updated this week
- Run LLaMA inference on CPU, with Rust 🦀🚀🦙☆35Jan 5, 2025Updated last year
- Rust native ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT2,...)☆3,060Jan 13, 2026Updated 4 months ago
- A WebGPU-accelerated ONNX inference run-time written 100% in Rust, ready for native and the web☆1,752Jul 21, 2024Updated last year
- Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference☆2,915Updated this week
- Fast, flexible LLM inference☆7,161Updated this week
- Rust bindings to https://github.com/ggerganov/whisper.cpp☆939Jul 30, 2025Updated 9 months ago
- Linear algebra foundation for the Rust programming language☆2,522May 3, 2026Updated 3 weeks ago
- Pure Rust implementation of a minimal Generative Pretrained Transformer☆935Oct 21, 2025Updated 7 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- High-level, optionally asynchronous Rust bindings to llama.cpp☆246Jun 5, 2024Updated last year
- ☆28Aug 10, 2023Updated 2 years ago
- Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.