A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
โ902Sep 1, 2026Updated this week
Alternatives and similar repositories for mosec
Users that are interested in mosec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ๐๏ธ Reproducible development environment for humans and agentsโ2,227Jul 25, 2026Updated last month
- An efficient binary serialization format for numerical data.โ18Nov 3, 2025Updated 10 months ago
- Autoscale LLM (vLLM, SGLang, LMDeploy) inferences on Kubernetes (and others)โ283Nov 3, 2023Updated 2 years ago
- This repository contains statistics about the AI Infrastructure products.โ16Feb 27, 2025Updated last year
- The Triton Inference Server provides an optimized cloud and edge inferencing solution.โ10,958Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer โข AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpaliโ2,929Mar 24, 2026Updated 5 months ago
- Fast, flexible LLM inferenceโ7,647Updated this week
- Large Language Model Text Generation Inferenceโ10,889Mar 21, 2026Updated 5 months ago
- MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.โ2,112Jun 30, 2025Updated last year
- Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kuberneteโฆโ2,233Updated this week
- Efficient, scalable and enterprise-grade CPU/GPU inference server for ๐ค Hugging Face transformer models ๐โ1,690Oct 23, 2024Updated last year
- OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)โ275Oct 11, 2023Updated 2 years ago
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.โ8,041Updated this week
- Large-scale model inference.โ628Sep 12, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI โข AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!โ8,818Updated this week
- Scalable, Low-latency and Hybrid-enabled Vector Search in Postgres. Revolutionize Vector Search, not Database.โ2,185Feb 26, 2025Updated last year
- Serving multiple LoRA finetuned LLM as oneโ1,175May 8, 2024Updated 2 years ago
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabiliโฆโ4,260Updated this week
- A conversational, AI device + software framework for companionship, entertainment, education, healthcare, IoT applications, and DIY robotโฆโ549Mar 27, 2026Updated 5 months ago
- Benchmark for machine learning model online serving (LLM, embedding, Stable-Diffusion, Whisper)โ28Jun 28, 2023Updated 3 years ago
- S-LoRA: Serving Thousands of Concurrent LoRA Adaptersโ1,923Jan 21, 2024Updated 2 years ago
- Transformer related optimization, including BERT, GPTโ6,448Mar 27, 2024Updated 2 years ago
- a fast cross platform AI inference engine ๐ค using Rust ๐ฆ and WebGPU ๐ฎโ470Jan 4, 2025Updated last year
- Managed Database hosting by DigitalOcean โข AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Turn PostgreSQL into your search engine in a Pythonic way.โ52Aug 29, 2025Updated last year
- AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (Nโฆโ4,725Aug 7, 2026Updated 3 weeks ago
- Training and serving large-scale neural networks with auto parallelization.โ3,178Dec 9, 2023Updated 2 years ago
- Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackabโฆโ1,584Jan 28, 2026Updated 7 months ago
- The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build cuโฆโ10,547Updated this week
- Accessible large language models via k-bit quantization for PyTorch.โ8,452Aug 27, 2026Updated last week
- Docker for Your ML/DL Models Based on OCI Artifactsโ474Jan 26, 2024Updated 2 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.โ33,326Updated this week
- torch::deploy (multipy for non-torch uses) is a system that lets you get around the GIL problem by running multiple Python interpreters iโฆโ179Dec 16, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer โข AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Running large language models on a single GPU for throughput-oriented scenarios.โ9,352Oct 28, 2024Updated last year
- Train transformer language models with reinforcement learning.โ19,198Updated this week
- An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifacโฆโ1,409Updated this week
- RayLLM - LLMs on Ray (Archived). Read README for more info.โ1,260Mar 13, 2025Updated last year
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.โ6,468Updated this week
- An awesome & curated list of best LLMOps tools for developersโ5,920May 21, 2026Updated 3 months ago
- All-in-one platform for search, recommendations, RAG, and analytics offered via APIโ2,716Jan 25, 2026Updated 7 months ago