EricLBuehler / candle-vllm
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
☆229Updated 3 weeks ago
Related projects: ⓘ
- Low rank adaptation (LoRA) for Candle.☆124Updated 3 weeks ago
- Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package☆138Updated 3 weeks ago
- Tutorial for Porting PyTorch Transformer Models to Candle (Rust)☆235Updated last month
- LLM Orchestrator built in Rust☆261Updated 6 months ago
- Library for generating vector embeddings, reranking in Rust☆250Updated 3 weeks ago
- Inference Llama 2 in one file of pure Rust 🦀☆227Updated last year
- High-level, optionally asynchronous Rust bindings to llama.cpp☆161Updated 3 months ago
- ☆134Updated this week
- LLama.cpp rust bindings☆320Updated 2 months ago
- ☆122Updated 4 months ago
- Llama2 LLM ported to Rust burn☆272Updated 5 months ago
- 🦀 A curated list of Rust tools, libraries, and frameworks for working with LLMs, GPT, AI☆236Updated 6 months ago
- OpenAI compatible API for serving LLAMA-2 model☆212Updated 10 months ago
- An LLM interface (chat bot) implemented in pure Rust using HuggingFace/Candle over Axum Websockets, an SQLite Database, and a Leptos (Was…☆118Updated last month
- Fast, streaming indexing and query library for AI (RAG) applications, written in Rust☆129Updated this week
- Extract core logic from qdrant and make it available as a library.☆56Updated 5 months ago
- Rust client for Qdrant vector search engine☆217Updated 3 weeks ago
- Rust wrapper for Microsoft's ONNX Runtime (version 1.8)☆275Updated 6 months ago
- A single-binary, GPU-accelerated LLM server (HTTP and WebSocket API) written in Rust☆79Updated 8 months ago
- Models and examples built with Burn☆162Updated this week
- Rust+OpenCL+AVX2 implementation of LLaMA inference code☆537Updated 7 months ago
- A Rust implementation of OpenAI's Whisper model using the burn framework☆262Updated 4 months ago
- LLaMa 7b with CUDA acceleration implemented in rust. Minimal GPU memory needed!☆100Updated last year
- ☆134Updated 7 months ago
- Rust implementation of the HNSW algorithm (Malkov-Yashunin)☆152Updated 2 months ago
- Example of tch-rs on M1☆47Updated 6 months ago
- 🦀Rust + Large Language Models - Make AI Services Freely and Easily.☆180Updated 6 months ago
- pgvector support for Rust☆114Updated last month
- Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (…☆289Updated 11 months ago
- ☆57Updated last year