Rust+OpenCL+AVX2 implementation of LLaMA inference code
☆554Feb 12, 2024Updated 2 years ago
Alternatives and similar repositories for rllama
Users that are interested in rllama are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A relatively basic implementation of RWKV in Rust written by someone with very little math and ML knowledge. Supports 32, 8 and 4 bit eva…☆95Sep 2, 2023Updated 2 years ago
- [Unmaintained, see README] An ecosystem of Rust libraries for working with large language models☆6,156Jun 24, 2024Updated 2 years ago
- Deep learning in Rust, with shape checked tensors and neural networks☆1,927Jul 23, 2024Updated 2 years ago
- Bleeding edge low level Rust binding for GGML☆18Jun 26, 2024Updated 2 years ago
- An implementation of the diffusers api in Rust☆595Apr 4, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A fast llama2 decoder in pure Rust.☆1,062Nov 30, 2023Updated 2 years ago
- `llm-chain` is a powerful rust crate for building chains in large language models allowing you to summarise text and complete complex tas…☆1,605Oct 31, 2024Updated last year
- LLaMa 7b with CUDA acceleration implemented in rust. Minimal GPU memory needed!☆112Jul 27, 2023Updated 3 years ago
- ☆58Apr 6, 2023Updated 3 years ago
- Inference Llama 2 in one file of pure Rust 🦀☆235Sep 11, 2023Updated 2 years ago
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.☆15,751Updated this week
- Rust bindings for the C++ api of PyTorch.☆5,474Jul 17, 2026Updated 3 weeks ago
- LLama.cpp rust bindings☆425Jun 27, 2024Updated 2 years ago
- Minimalist ML framework for Rust☆20,896Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Llama2 LLM ported to Rust burn☆279Apr 16, 2024Updated 2 years ago
- Run LLaMA inference on CPU, with Rust 🦀🚀🦙☆35Jan 5, 2025Updated last year
- Rust native ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT2,...)☆3,076Jan 13, 2026Updated 7 months ago
- A WebGPU-accelerated ONNX inference run-time written 100% in Rust, ready for native and the web☆1,754Jul 21, 2024Updated 2 years ago
- Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference☆3,029Updated this week
- Rust bindings to https://github.com/ggerganov/whisper.cpp☆947Jul 30, 2025Updated last year
- Fast, flexible LLM inference☆7,591Jul 29, 2026Updated 2 weeks ago
- Pure Rust implementation of a minimal Generative Pretrained Transformer☆936Oct 21, 2025Updated 9 months ago
- Linear algebra foundation for the Rust programming language☆2,565Jun 24, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- High-level, optionally asynchronous Rust bindings to llama.cpp☆252Jun 5, 2024Updated 2 years ago
- ☆28Aug 10, 2023Updated 3 years ago
- Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.☆5,310Apr 29, 2026Updated 3 months ago
- Tensors and dynamic neural networks in pure Rust.☆1,086Oct 10, 2022Updated 3 years ago
- Safe rust wrapper around CUDA toolkit☆1,205Updated this week
- A Discord bot, written in Rust, that generates responses using the LLaMA language model.☆94Aug 13, 2023Updated 3 years ago
- ndarray: an N-dimensional array with array views, multidimensional slicing, and efficient operations☆4,312Jul 18, 2026Updated 3 weeks ago
- A Rust machine learning framework.☆4,728May 30, 2026Updated 2 months ago
- ☆13Aug 29, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- C++ implementation for BLOOM☆811May 13, 2023Updated 3 years ago
- Instant, controllable, local pre-trained AI models in Rust☆2,218Updated this week
- Fast ML inference & training for ONNX models in Rust☆2,446Updated this week
- A Rust implementation of the Khronos OpenCL 3.0 API.☆134Mar 8, 2026Updated 5 months ago
- ggml implementation of BERT☆502Feb 23, 2024Updated 2 years ago
- Port of Microsoft's BioGPT in C/C++ using ggml☆87Feb 21, 2024Updated 2 years ago
- SoTA Transformers with C-backend for fast inference on your CPU.☆312Dec 9, 2023Updated 2 years ago