octoml / triton-client-rsLinks
A client library in Rust for Nvidia Triton.
☆30Updated last year
Alternatives and similar repositories for triton-client-rs
Users that are interested in triton-client-rs are comparing it to the libraries listed below
Sorting:
- Rust wrapper for Microsoft's ONNX Runtime (version 1.8)☆297Updated last year
- Rust crate for some audio utilities☆24Updated 3 months ago
- ☆30Updated 7 months ago
- Example of tch-rs on M1☆53Updated last year
- Asynchronous CUDA for Rust.☆33Updated 7 months ago
- Your one stop CLI for ONNX model analysis.☆47Updated 2 years ago
- Low rank adaptation (LoRA) for Candle.☆150Updated 2 months ago
- ☆23Updated 2 months ago
- Dataflow is a data processing library, primarily for machine learning.☆22Updated 2 years ago
- A collection of optimisers for use with candle☆36Updated last month
- Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package☆209Updated last week
- ☆26Updated last year
- Rust library for whisper.cpp compatible Mel spectrograms☆70Updated last month
- Rust wrapper for Microsoft's ONNX Runtime with CUDA support (version 1.7)☆23Updated 2 years ago
- Proof of concept for running moshi/hibiki using webrtc☆19Updated 3 months ago
- Fast and versatile tokenizer for language models, compatible with SentencePiece, Tokenizers, Tiktoken and more. Supports BPE, Unigram and…☆26Updated 3 months ago
- Rust library for running TensorRT accelerated deep learning models☆56Updated 3 years ago
- An extension library to Candle that provides PyTorch functions not currently available in Candle☆39Updated last year
- GPU based FFT written in Rust and CubeCL☆23Updated 2 weeks ago
- implement llava using candle☆15Updated last year
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆380Updated this week
- Rust port of sentence-transformers (https://github.com/UKPLab/sentence-transformers)☆115Updated 9 months ago
- Modular Rust transformer/LLM library using Candle☆36Updated last year
- Fast serverless LLM inference, in Rust.☆77Updated 3 months ago
- Experimental ONNX implementation for WASI NN.☆48Updated 3 years ago
- Extract core logic from qdrant and make it available as a library.☆59Updated last year
- A complete(grpc service and lib) Rust inference with multilingual embedding support. This version leverages the power of Rust for both GR…☆39Updated 10 months ago
- Experimental compiler for deep learning models☆67Updated last month
- ☆39Updated 2 years ago
- ☆20Updated 8 months ago