just-joseph / llmops-realtime-inference-pipelineView on GitHub
Deploy Large Language Models for real-time inference at scale. Complete MLOps pipeline: BentoML + Docker + Kubernetes + AWS EKS + GitHub Actions + Prometheus/Grafana monitoring. Demonstrates production LLM serving with auto-scaling, load balancing, and performance optimization.
22Jan 18, 2026Updated 6 months ago

Alternatives and similar repositories for llmops-realtime-inference-pipeline

Users that are interested in llmops-realtime-inference-pipeline are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?