Self-host LLMs with LMDeploy and BentoML
☆22Jul 14, 2026Updated 3 weeks ago
Alternatives and similar repositories for BentoLMDeploy
Users that are interested in BentoLMDeploy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Stateful LLM Serving☆105Mar 11, 2025Updated last year
- Deepseek-CoT☆10Oct 6, 2024Updated last year
- ☆17Jan 31, 2026Updated 6 months ago
- Prompt templates for language models☆10Apr 7, 2026Updated 4 months ago
- A curated list for Efficient Large Language Models☆11Mar 25, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Langchain + Docker + Neo4j☆10Oct 29, 2024Updated last year
- Details of the datasets for Few-shot class-incremental audio classification☆11Dec 6, 2023Updated 2 years ago
- ☆20Apr 24, 2023Updated 3 years ago
- letta integration for terminalbench (#1 open source agent, in under 200 lines of code)☆19Oct 22, 2025Updated 9 months ago
- API serving for your diffusers models☆11Jan 19, 2024Updated 2 years ago
- Final training script from HuggingFace Whisper Fine tuning event - to get best results on finetuned model.☆12Dec 24, 2022Updated 3 years ago
- Sentence Embedding as a Service☆15Jun 30, 2025Updated last year
- Interface Design for Self-Supervised Speech Models, Accepted to Interspeech2024☆16Nov 19, 2024Updated last year
- Distributed IO-aware Attention algorithm☆24Sep 24, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆13Oct 27, 2021Updated 4 years ago
- Elevate your language models with insightful diversity metrics.☆11Feb 4, 2024Updated 2 years ago
- MUX-PLMs: Pretraining LMs with Data Multiplexing☆15Jan 29, 2023Updated 3 years ago
- A RAG that can scale 🧑🏻💻☆11May 28, 2024Updated 2 years ago
- ☆14May 12, 2025Updated last year
- A throughput-oriented high-performance serving framework for LLMs☆974Mar 29, 2026Updated 4 months ago
- ☆11Jul 16, 2026Updated 3 weeks ago
- Simple Telegram bot to annotate and varify automatic speech recognition datasets☆12Mar 30, 2021Updated 5 years ago
- Demonstrate Function Calling code portability across 4 AI Models: OpenAI, AzureOpenAI, VertexAI Gemini and Mistral AI.☆13Jun 7, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Supplemental materials for The ASPLOS 2025 / EuroSys 2025 Contest on Intra-Operator Parallelism for Distributed Deep Learning☆25May 12, 2025Updated last year
- A fast parallel implementation of RNN Transducer.☆11Apr 8, 2025Updated last year
- 3D visualization of depth maps which are created by any depth model such as Monodepth, Packnet, etc..☆12Jun 9, 2020Updated 6 years ago
- Resources related to EMNLP 2021 paper "FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input Representations"☆13Dec 14, 2021Updated 4 years ago
- Source code for ACL 2020 paper "Learning Spoken Language Representations with Neural Lattice Language Modeling"☆18Feb 11, 2023Updated 3 years ago
- ROUGE L metric implementation using tensorflow ops☆12Sep 17, 2018Updated 7 years ago
- ☆11May 18, 2025Updated last year
- EuroSys '24: "Trinity: A Fast Compressed Multi-attribute Data Store"☆18Mar 8, 2025Updated last year
- ☆16Nov 8, 2020Updated 5 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Fast model deployment on AWS EC2☆14Feb 25, 2024Updated 2 years ago
- [APSIPA'22] Exploring Speaker Age Estimation on Different Self-Supervised Learning Models☆14Oct 19, 2022Updated 3 years ago
- [ICLR 2025] SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-Correction☆91Mar 23, 2025Updated last year
- Code of the paper "Low-Latency Speech Separation Guided Diarization for Telephone Conversations"☆15Dec 22, 2022Updated 3 years ago
- Implementation of SmoothCache, a project aimed at speeding-up Diffusion Transformer (DiT) based GenAI models with error-guided caching.☆48Jul 17, 2025Updated last year
- ☆15Jun 10, 2024Updated 2 years ago
- ☆14Jul 14, 2026Updated 3 weeks ago