octoml / triton-client-rsLinks
A client library in Rust for Nvidia Triton.
☆30Updated 2 years ago
Alternatives and similar repositories for triton-client-rs
Users that are interested in triton-client-rs are comparing it to the libraries listed below
Sorting:
- Rust wrapper for Microsoft's ONNX Runtime (version 1.8)☆308Updated last year
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆502Updated last week
- Rust crate for some audio utilities☆25Updated 7 months ago
- Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package☆236Updated last month
- ☆24Updated 6 months ago
- Your one stop CLI for ONNX model analysis.☆47Updated 2 years ago
- Example of tch-rs on M1☆55Updated last year
- Fast and versatile tokenizer for language models, compatible with SentencePiece, Tokenizers, Tiktoken and more. Supports BPE, Unigram and…☆36Updated 2 weeks ago
- A collection of optimisers for use with candle☆43Updated 2 months ago
- Low rank adaptation (LoRA) for Candle.☆163Updated 6 months ago
- Asynchronous CUDA for Rust.☆36Updated last month
- Rust library for whisper.cpp compatible Mel spectrograms☆75Updated 5 months ago
- Dataflow is a data processing library, primarily for machine learning.☆24Updated 2 years ago
- A complete(grpc service and lib) Rust inference with multilingual embedding support. This version leverages the power of Rust for both GR…☆39Updated last year
- Rust port of sentence-transformers (https://github.com/UKPLab/sentence-transformers)☆121Updated last year
- Fast serverless LLM inference, in Rust.☆94Updated 7 months ago
- An extension library to Candle that provides PyTorch functions not currently available in Candle☆40Updated last year
- Automatically derive Python dunder methods for your Rust code☆20Updated 6 months ago
- Rust language bindings for Faiss☆239Updated last month
- ☆36Updated 11 months ago
- Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (…☆328Updated 2 years ago
- A single-binary, GPU-accelerated LLM server (HTTP and WebSocket API) written in Rust☆79Updated last year
- An example of using Torch rust bindings to serve trained machine learning models via Actix Web☆16Updated 4 years ago
- CLI utility to inspect and explore .safetensors and .gguf files☆32Updated 3 months ago
- Tutorial for Porting PyTorch Transformer Models to Candle (Rust)☆325Updated last year
- LLaMa 7b with CUDA acceleration implemented in rust. Minimal GPU memory needed!☆110Updated 2 years ago
- implement llava using candle☆15Updated last year
- Tantivy directory implementation backed by object_store☆36Updated last year
- Sample Python extension using Rust/PyO3/tch to interact with PyTorch☆37Updated last year
- GPU based FFT written in Rust and CubeCL☆24Updated 4 months ago