llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more.
☆95May 17, 2024Updated 2 years ago
Alternatives and similar repositories for llm-inference
Users that are interested in llm-inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The framework of training large language models,support lora, full parameters fine tune etc, define yaml to start training/fine tune of y…☆31Sep 19, 2024Updated last year
- The CSGHub SDK is a powerful Python client specifically designed to interact seamlessly with the CSGHub server. This toolkit is engineere…☆21Jul 17, 2026Updated last week
- LLM scheduler user interface☆21May 17, 2024Updated 2 years ago
- AutoHub: A Personal Browser Automation Assistant☆25Jul 30, 2025Updated 11 months ago
- ☆13Jan 7, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆20Sep 28, 2024Updated last year
- A side project that follows all the acceleration tricks in tinyllama, with the minimal modification to the huggingface transformers code.☆13Sep 2, 2024Updated last year
- Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).☆250Mar 15, 2024Updated 2 years ago
- ☆11Jan 8, 2025Updated last year
- 3D Multi-person Pose Estimation in Multi-view Environment using 3D U-Net Transformer Networks☆21Jun 7, 2021Updated 5 years ago
- Codes for our paper "AgentMonitor: A Plug-and-Play Framework for Predictive and Secure Multi-Agent Systems"☆13Dec 13, 2024Updated last year
- Large Language Model Onnx Inference Framework☆35Nov 25, 2025Updated 8 months ago
- Secure and Scalable Federated Learning using Serverless Computing☆13Jan 31, 2024Updated 2 years ago
- Device plugins for Volcano, e.g. GPU☆137Mar 20, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Finance SaaS Platform with ability to track your income and expenses, categorize transactions and assign them to specific accounts, as w…☆16Dec 20, 2024Updated last year
- NVIDIA TensorRT Hackathon 2023复赛选题:通义千问Qwen-7B用TensorRT-LLM模型搭建及优化☆43Oct 20, 2023Updated 2 years ago
- 官方transformers源码解析。AI大模型时代,pytorch、transformer是新操作系统,其他都是运行在其上面的软件。☆16Sep 25, 2023Updated 2 years ago
- An intelligent chat interface designed for seamless interaction with Model Context Protocol (MCP) tools.☆19May 26, 2025Updated last year
- WebRTC-based real-time audio streaming with Faster Whisper ASR integration for live speech-to-text transcription.☆13Sep 27, 2024Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMs☆109Updated this week
- Deduplication over dis-aggregated memory for Serverless Computing☆14Mar 21, 2022Updated 4 years ago
- A unified programming framework for high and portable performance across FPGAs and GPUs☆11Mar 23, 2025Updated last year
- DISB is a new DNN inference serving benchmark with diverse workloads and models, as well as real-world traces.☆58Aug 21, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆168Oct 9, 2024Updated last year
- 基于ollama推理框架本地部署的Agent应用,实现MCP工具调用,短长期记忆等功能。|| A locally deployed agent application built on the Ollama, featuring MCP tool integration …☆16Dec 10, 2025Updated 7 months ago
- Mathematical expression evaluator with just in time code generation.☆12Apr 7, 2013Updated 13 years ago
- Pretrain, finetune and serve LLMs on Intel platforms with Ray☆130Sep 23, 2025Updated 10 months ago
- Selection-based Question Answering☆14Feb 7, 2018Updated 8 years ago
- Distributed SDDMM Kernel☆12Jul 8, 2022Updated 4 years ago
- ☢️ TensorRT 2023复赛——基于TensorRT-LLM的Llama模型推断加速优化☆54Oct 20, 2023Updated 2 years ago
- Opinionated Langchain setup with Qdrant vector store and Kong gateway☆32Apr 7, 2023Updated 3 years ago
- ☆14Apr 23, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Neural Memory Compression System for RAG Applications☆20Nov 20, 2025Updated 8 months ago
- ☆25Aug 27, 2021Updated 4 years ago
- CUDA C simple application for Nvidia's GPU☆11Jun 7, 2022Updated 4 years ago
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 2 months ago
- ☆12Mar 31, 2021Updated 5 years ago
- ☆31May 13, 2024Updated 2 years ago
- Fork of NACA from Google Code☆13Feb 25, 2010Updated 16 years ago