Triton CLI is an open source command line interface that enables users to create, deploy, and profile models served by the Triton Inference Server.
☆73Aug 7, 2026Updated last week
Alternatives and similar repositories for triton_cli
Users that are interested in triton_cli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Integrating SSE with NVIDIA Triton Inference Server using a Python backend and Zephyr model. There is very less documentation how to use …☆10May 29, 2024Updated 2 years ago
- Triton Model Navigator is an inference toolkit designed for optimizing and deploying Deep Learning models with a focus on NVIDIA GPUs.☆224May 27, 2026Updated 2 months ago
- Triton backend that enables pre-process, post-processing and other logic to be implemented in Python.☆681Updated this week
- An api for interfacing Nvidia Trition Inference Server with Rust☆12Jun 12, 2023Updated 3 years ago
- Repository for open inference protocol specification☆77May 12, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository contains tutorials and examples for Triton Inference Server☆858Aug 7, 2026Updated last week
- TRITONCACHE implementation of a Redis cache☆18Aug 7, 2026Updated last week
- Triton Model Analyzer is a CLI tool to help with better understanding of the compute and memory requirements of the Triton Inference Serv…☆523Aug 7, 2026Updated last week
- Please visit https://github.com/HKUSTDial/NL2SQL360 to get the official code!☆10Sep 1, 2024Updated last year
- ☆153Aug 7, 2026Updated last week
- Code for "Moving on from OntoNotes: Coreference Resolution Model Transfer" and "Incremental Neural Coreference Resolution in Constant Mem…☆18Mar 11, 2022Updated 4 years ago
- PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.☆847Aug 13, 2025Updated last year
- VibeRL is a Reinforcement Learning framework built essentially through vibe coding with Kimi K2.☆17Updated this week
- The Triton TensorRT-LLM Backend☆941Jul 22, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Quotek is an open source algotrading platform, written in C++.☆10Nov 12, 2020Updated 5 years ago
- Simple python library for generating your own perfetto traces for your application. Can be used for both app instrumentation and custom …☆27Jun 22, 2025Updated last year
- Multiple GEMM operators are constructed with cutlass to support LLM inference.☆20Aug 3, 2025Updated last year
- Multi-layer perceptron, Autoencoder, and Restricted Boltzmann Machine☆10Sep 15, 2018Updated 7 years ago
- This repository provides optical character detection and recognition solution optimized on Nvidia devices.☆90May 13, 2025Updated last year
- ☆22Aug 7, 2026Updated last week
- The core library and APIs implementing the Triton Inference Server.☆179Updated this week
- 专注于把fastgpt对接到微信生态中,此项目主要贡献给小胰宝、小肺宝等系列公益项目使用☆19Mar 9, 2025Updated last year
- An Operator for deployment and maintenance of NVIDIA NIMs and NeMo microservices in a Kubernetes environment.☆160Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- An NVIDIA Triton Server workflow for OCR and the LayoutLMv3 Transformer Model☆30Sep 14, 2022Updated 3 years ago
- Compare multiple optimization methods on triton to imporve model service performance☆52Jan 10, 2024Updated 2 years ago
- Code for the EMNLP 2020 paper "Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks"☆25Jul 3, 2021Updated 5 years ago
- An NVIDIA AI Workbench example project to build a multimodal virtual assistant☆22Apr 17, 2025Updated last year
- A library for enabling databinding to and from Observables in xaml☆10Jan 6, 2018Updated 8 years ago
- MIG Partition Editor for NVIDIA GPUs☆260Updated this week
- Vast-ai public repository for open sourced tools, plugins, etc.☆17Nov 4, 2024Updated last year
- Implementation of SmoothCache, a project aimed at speeding-up Diffusion Transformer (DiT) based GenAI models with error-guided caching.☆48Jul 17, 2025Updated last year
- Support for language highlighting of KECC(KAIST Educational C Compiler) IR☆12May 17, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- SGLang Kernel Wheel Index☆25Updated this week
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆321Updated this week
- A brief understanding of ffmpeg cli through pseudocode☆11Dec 20, 2020Updated 5 years ago
- The Triton backend for TensorFlow.☆57Nov 22, 2025Updated 8 months ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated 11 months ago
- Mid-Level Rust Bindings to the C API for Microsoft's ONNX Runtime☆18Aug 21, 2024Updated last year
- ☆352Aug 7, 2026Updated last week