llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more.
☆97May 17, 2024Updated 2 years ago
Alternatives and similar repositories for llm-inference
Users that are interested in llm-inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The framework of training large language models,support lora, full parameters fine tune etc, define yaml to start training/fine tune of y…☆32Sep 19, 2024Updated 2 years ago
- The CSGHub SDK is a powerful Python client specifically designed to interact seamlessly with the CSGHub server. This toolkit is engineere…☆24Aug 29, 2026Updated 3 weeks ago
- LLM scheduler user interface☆21May 17, 2024Updated 2 years ago
- AutoHub: A Personal Browser Automation Assistant☆25Jul 30, 2025Updated last year
- ☆17Mar 24, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆13Jan 7, 2025Updated last year
- ☆20Sep 28, 2024Updated last year
- A side project that follows all the acceleration tricks in tinyllama, with the minimal modification to the huggingface transformers code.☆13Sep 2, 2024Updated 2 years ago
- RayLLM - LLMs on Ray (Archived). Read README for more info.☆1,261Mar 13, 2025Updated last year
- Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).☆251Mar 15, 2024Updated 2 years ago
- ☆11Jan 8, 2025Updated last year
- Elastic-Grok-Script-Plugin is a provider of Grok ElasticSearch plug-in☆12Dec 6, 2016Updated 9 years ago
- Codes for our paper "AgentMonitor: A Plug-and-Play Framework for Predictive and Secure Multi-Agent Systems"☆14Dec 13, 2024Updated last year
- Large Language Model Onnx Inference Framework☆35Nov 25, 2025Updated 10 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Secure and Scalable Federated Learning using Serverless Computing☆14Jan 31, 2024Updated 2 years ago
- ☆75Mar 26, 2025Updated last year
- Azure Machine Learning - MLOps Python SDKv2☆10Jul 24, 2023Updated 3 years ago
- WebRTC-based real-time audio streaming with Faster Whisper ASR integration for live speech-to-text transcription.☆13Sep 27, 2024Updated 2 years ago
- WIP. Veloce is a low-code Ray-based parallelization library that makes machine learning computation novel, efficient, and heterogeneous.☆17Aug 4, 2022Updated 4 years ago
- Deduplication over dis-aggregated memory for Serverless Computing☆14Mar 21, 2022Updated 4 years ago
- A unified programming framework for high and portable performance across FPGAs and GPUs☆11Mar 23, 2025Updated last year
- This repository contains the results and code for the MLPerf™ Inference v2.1 benchmark.☆18Jul 24, 2025Updated last year
- Dynamic Memory Management for Serving LLMs without PagedAttention☆524Aug 24, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- DISB is a new DNN inference serving benchmark with diverse workloads and models, as well as real-world traces.☆58Aug 21, 2024Updated 2 years ago
- ☆172Oct 9, 2024Updated last year
- Kubernetes device plugin for Biren GPU☆11Oct 17, 2024Updated last year
- An adaption of Senders/Receivers for async networking and I/O☆21Apr 25, 2025Updated last year
- Pretrain, finetune and serve LLMs on Intel platforms with Ray☆129Sep 23, 2025Updated last year
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- 日志采集工具☆20Aug 16, 2017Updated 9 years ago
- ☢️ TensorRT 2023复赛——基于TensorRT-LLM的Llama模型推断加速优化☆54Oct 20, 2023Updated 2 years ago
- [ICLR 2025] A trinity of environments, tools, and benchmarks for general virtual agents☆233Jun 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Light local website for displaying performances from different chat models.☆86Nov 13, 2023Updated 2 years ago
- Neural Memory Compression System for RAG Applications☆21Nov 20, 2025Updated 10 months ago
- Compare different hardware platforms via the Roofline Model for LLM inference tasks.☆126Mar 13, 2024Updated 2 years ago
- ☆12Mar 31, 2021Updated 5 years ago
- ☆31May 13, 2024Updated 2 years ago
- Console application and Go package (library) to validate semantic versions☆10Mar 22, 2021Updated 5 years ago
- Fork of NACA from Google Code☆13Feb 25, 2010Updated 16 years ago