A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
โ904Jul 11, 2026Updated 2 weeks ago
Alternatives and similar repositories for mosec
Users that are interested in mosec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ๐๏ธ Reproducible development environment for humans and agentsโ2,214Jul 10, 2026Updated 2 weeks ago
- An efficient binary serialization format for numerical data.โ18Nov 3, 2025Updated 8 months ago
- Autoscale LLM (vLLM, SGLang, LMDeploy) inferences on Kubernetes (and others)โ282Nov 3, 2023Updated 2 years ago
- This repository contains statistics about the AI Infrastructure products.โ16Feb 27, 2025Updated last year
- The Triton Inference Server provides an optimized cloud and edge inferencing solution.โ10,868Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer โข AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpaliโ2,892Mar 24, 2026Updated 4 months ago
- Fast, flexible LLM inferenceโ7,518Updated this week
- Large Language Model Text Generation Inferenceโ10,882Mar 21, 2026Updated 4 months ago
- MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.โ2,108Jun 30, 2025Updated last year
- Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kuberneteโฆโ2,192Updated this week
- Efficient, scalable and enterprise-grade CPU/GPU inference server for ๐ค Hugging Face transformer models ๐โ1,690Oct 23, 2024Updated last year
- OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)โ276Oct 11, 2023Updated 2 years ago
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.โ7,972Updated this week
- Large-scale model inference.โ629Sep 12, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off โข AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!โ8,734Updated this week
- Scalable, Low-latency and Hybrid-enabled Vector Search in Postgres. Revolutionize Vector Search, not Database.โ2,180Feb 26, 2025Updated last year
- Serving multiple LoRA finetuned LLM as oneโ1,167May 8, 2024Updated 2 years ago
- LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalabiliโฆโ4,193Updated this week
- A conversational, AI device + software framework for companionship, entertainment, education, healthcare, IoT applications, and DIY robotโฆโ549Mar 27, 2026Updated 3 months ago
- Benchmark for machine learning model online serving (LLM, embedding, Stable-Diffusion, Whisper)โ28Jun 28, 2023Updated 3 years ago
- S-LoRA: Serving Thousands of Concurrent LoRA Adaptersโ1,920Jan 21, 2024Updated 2 years ago
- Transformer related optimization, including BERT, GPTโ6,444Mar 27, 2024Updated 2 years ago
- a fast cross platform AI inference engine ๐ค using Rust ๐ฆ and WebGPU ๐ฎโ470Jan 4, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer โข AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Turn PostgreSQL into your search engine in a Pythonic way.โ51Aug 29, 2025Updated 10 months ago
- AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (Nโฆโ4,726Jul 14, 2026Updated last week
- Training and serving large-scale neural networks with auto parallelization.โ3,180Dec 9, 2023Updated 2 years ago
- Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackabโฆโ1,585Jan 28, 2026Updated 5 months ago
- The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build cuโฆโ10,382Updated this week
- Accessible large language models via k-bit quantization for PyTorch.โ8,338Updated this week
- Docker for Your ML/DL Models Based on OCI Artifactsโ473Jan 26, 2024Updated 2 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.โ30,706Updated this week
- torch::deploy (multipy for non-torch uses) is a system that lets you get around the GIL problem by running multiple Python interpreters iโฆโ179Dec 16, 2025Updated 7 months ago
- Managed Database hosting by DigitalOcean โข AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Train transformer language models with reinforcement learning.โ18,920Updated this week
- Running large language models on a single GPU for throughput-oriented scenarios.โ9,363Oct 28, 2024Updated last year
- An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifacโฆโ1,390Updated this week
- RayLLM - LLMs on Ray (Archived). Read README for more info.โ1,262Mar 13, 2025Updated last year
- An awesome & curated list of best LLMOps tools for developersโ5,894May 21, 2026Updated 2 months ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.โ5,986Updated this week
- All-in-one platform for search, recommendations, RAG, and analytics offered via APIโ2,698Jan 25, 2026Updated 5 months ago