Lightweight continuous batching OpenAI compatibility using HuggingFace Transformers include T5 and Whisper.
☆29Mar 15, 2025Updated last year
Alternatives and similar repositories for transformers-continuous-batching
Users that are interested in transformers-continuous-batching are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆11Feb 20, 2025Updated last year
- An fully autonomous agent that accesses the browser and performs tasks.☆18Apr 25, 2025Updated last year
- A MCP stdio toolpack for local LLMs☆33Apr 6, 2026Updated 3 months ago
- This is a FastAPI based LLM server. Load multiple LLM models (MLX or llama.cpp) simultaneously using multiprocessing.☆18Apr 8, 2026Updated 3 months ago
- an auto-sleeping and -waking framework around llama.cpp☆13Feb 8, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆19Updated this week
- SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversati…☆29May 6, 2025Updated last year
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆25Apr 1, 2025Updated last year
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- Multi-turn dataset management tool for LLM trainers☆13Mar 31, 2025Updated last year
- Utility to use eleven lab's streaming to in the command line☆11Aug 8, 2023Updated 2 years ago
- implementation of https://arxiv.org/pdf/2312.09299☆21Jul 3, 2024Updated 2 years ago
- ☆24Jan 22, 2025Updated last year
- ☆16Dec 16, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Type to talk on Discord, Skype, and other voice chatting programs.☆10Jul 20, 2020Updated 6 years ago
- ☆18Sep 4, 2024Updated last year
- ☆40Mar 25, 2023Updated 3 years ago
- A custom LiteLLM provider enabling local execution of Hugging Face models with streaming, quantization, and async support☆30Jun 22, 2025Updated last year
- ☆14Mar 2, 2025Updated last year
- A Windows tool to query various LLM AIs. Supports branched conversations, history and summaries among others.☆36May 11, 2026Updated 2 months ago
- CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval☆13Jun 27, 2025Updated last year
- Replayable Browser Agent☆16Apr 24, 2026Updated 3 months ago
- A pytorch implementation of a text to videos GAN☆12Jul 26, 2019Updated 6 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆11Oct 11, 2023Updated 2 years ago
- Run Orpheus 3B Locally With LM Studio☆32Mar 20, 2025Updated last year
- Orpheus Chat WebUI☆76Mar 27, 2025Updated last year
- Demonstration that finetuning RoPE model on larger sequences than the pre-trained model adapts the model context limit☆62Jun 21, 2023Updated 3 years ago
- input aspect ratio, output dimensions☆21Mar 13, 2026Updated 4 months ago
- Recursive Language Model MCP Server - Enable any LLM to process arbitrarily long contexts☆15Jan 7, 2026Updated 6 months ago
- An extension of MCP for SillyTavern.☆89Jun 9, 2026Updated last month
- Automated Identification of Redundant Layer Blocks for Pruning in Large Language Models☆267Apr 23, 2024Updated 2 years ago
- Simulates talk with an AI that can express emotions☆89Apr 4, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A sleek, customizable interface for managing LLMs with responsive design and easy agent personalization.☆19Aug 30, 2024Updated last year
- Telegram dev bot for all your dirty work☆15Feb 9, 2026Updated 5 months ago
- Examples flows for Kestra☆15Jul 15, 2026Updated last week
- Evaluate state-of-the-art sparse embedding models on the LIMIT dataset (`limit-small` and `limit`) from google's paper `On the Theoretica…☆16Sep 4, 2025Updated 10 months ago
- Create embeddings for LLM using the Nomic API☆23Nov 21, 2024Updated last year
- Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient☆68Aug 3, 2025Updated 11 months ago
- 🎙️ VibeVoice FastAPI - Multi-Speaker TTS API☆32Aug 27, 2025Updated 10 months ago