octoml / triton-client-rsLinks
A client library in Rust for Nvidia Triton.
☆30Updated last year
Alternatives and similar repositories for triton-client-rs
Users that are interested in triton-client-rs are comparing it to the libraries listed below
Sorting:
- Rust wrapper for Microsoft's ONNX Runtime (version 1.8)☆301Updated last year
- Rust crate for some audio utilities☆26Updated 4 months ago
- ☆23Updated 3 months ago
- Example of tch-rs on M1☆54Updated last year
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆394Updated last week
- Your one stop CLI for ONNX model analysis.☆47Updated 2 years ago
- Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package☆214Updated last month
- Rust library for whisper.cpp compatible Mel spectrograms☆72Updated 2 months ago
- Fast serverless LLM inference, in Rust.☆88Updated 4 months ago
- ☆30Updated 7 months ago
- Low rank adaptation (LoRA) for Candle.☆152Updated 2 months ago
- implement llava using candle☆15Updated last year
- A collection of optimisers for use with candle☆36Updated last month
- LLaMa 7b with CUDA acceleration implemented in rust. Minimal GPU memory needed!☆108Updated last year
- Asynchronous CUDA for Rust.☆34Updated 8 months ago
- Dataflow is a data processing library, primarily for machine learning.☆23Updated 2 years ago
- A complete(grpc service and lib) Rust inference with multilingual embedding support. This version leverages the power of Rust for both GR…☆39Updated 10 months ago
- Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (…☆321Updated last year
- A Fish Speech implementation in Rust, with Candle.rs☆94Updated last month
- GPU based FFT written in Rust and CubeCL☆23Updated last month
- A high-performance constrained decoding engine based on context free grammar in Rust☆54Updated last month
- An in-process trace collector using the Rust tracing framework and the Perfetto C++ SDK☆13Updated 3 months ago
- A Demo server serving Bert through ONNX with GPU written in Rust with <3☆40Updated 3 years ago
- ☆89Updated 6 months ago
- A single-binary, GPU-accelerated LLM server (HTTP and WebSocket API) written in Rust☆80Updated last year
- ☆39Updated 2 years ago
- The Triton backend for the PyTorch TorchScript models.☆156Updated last week
- Sample Python extension using Rust/PyO3/tch to interact with PyTorch☆37Updated last year
- An extension library to Candle that provides PyTorch functions not currently available in Candle☆40Updated last year
- ESRGAN implemented in rust with candle☆17Updated last year