A high performance batching router optimises max throughput for text inference workload
☆16Sep 6, 2023Updated 2 years ago
Alternatives and similar repositories for text-inference-batcher
Users that are interested in text-inference-batcher are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Various files of sometimes useful code snippits to use with Ollama☆10Jan 27, 2025Updated last year
- Using modal.com to process FineWeb-edu data☆20Apr 11, 2026Updated 4 months ago
- Get deterministic output in any format like json from any LLM.☆19Apr 25, 2023Updated 3 years ago
- the small distributed language model toolkit; fine-tune state-of-the-art LLMs anywhere, rapidly☆33Oct 19, 2024Updated last year
- 🚀 LLM inference optimization simulator, modeling compute-bound prefill and memory-bound decode phases.☆13Jul 12, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 11 months ago
- This library supports evaluating disparities in generated image quality, diversity, and consistency between geographic regions.☆20Jun 3, 2024Updated 2 years ago
- Unleash the power of FOSS language models on your local machine☆30Nov 21, 2023Updated 2 years ago
- Sparse autoencoders for Contra text embedding models☆25Apr 24, 2024Updated 2 years ago
- B-Llama3o a llama3 with Vision Audio and Audio understanding as well as text and Audio and Animation Data output.☆26Jun 3, 2024Updated 2 years ago
- Planning Poker built with Cloudflare Workers, Workers KV, Durable Objects, Websockets, and Cloudflare Pages. Also React, Redux Toolkit, T…☆19Oct 12, 2023Updated 2 years ago
- A Scheduler for Batched LLM Inference☆19Oct 5, 2025Updated 10 months ago
- Simple Tool Caller for llama.cpp☆11Aug 12, 2024Updated 2 years ago
- PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation☆32Nov 16, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Geospatial Next.js app with DuckDB-Wasm☆15May 24, 2023Updated 3 years ago
- ☆14Oct 31, 2023Updated 2 years ago
- You can use it to modify HTTP (S) response values, redirect static file requests to the local file directory, and support batch modificat…☆18Nov 30, 2022Updated 3 years ago
- Lightweight JavaScript library for parsing and manipulating TAK messages, primarily Cursor-on-Target (COT)☆13Mar 30, 2026Updated 4 months ago
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆24May 5, 2026Updated 3 months ago
- Diffusion_TTS extension for booga☆71Sep 6, 2025Updated 11 months ago
- A simple speech-to-text and text-to-speech AI chatbot that can be run fully offline.☆45Jan 28, 2024Updated 2 years ago
- Pynocular is a lightweight ORM that lets you query your database using Pydantic models and asyncio☆11May 24, 2022Updated 4 years ago
- Synthèses vocale piper oobabooga☆14Feb 24, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Typed REST queries - build, transport, validate, execute. @rapiq/* packages implement a JSON-API-style query scheme (fields, filters, rel…☆32Updated this week
- OCI images for the curious.☆12Mar 21, 2022Updated 4 years ago
- Debugger for HTC phones bootloader (HBOOT).☆20Nov 28, 2013Updated 12 years ago
- utilities for loading and running text embeddings with onnx☆46Aug 16, 2025Updated 11 months ago
- Blue-text Bot AI. Uses Ollama + AppleScript☆50May 19, 2024Updated 2 years ago
- Guide to deploying Slurm and OpenMPI on Raspberry Pi computers☆13Nov 9, 2024Updated last year
- Tiny Go libraries that I keep copying between projects☆12Updated this week
- Sync Taskwarrior tasks with Habitica☆15Sep 14, 2022Updated 3 years ago
- KPI Reporter is a dev-friendly, on-premises tool for crafting automated reports on what matters to you.☆10Oct 6, 2022Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- API gateway and reverse proxy for OpenAI APIs☆15Jul 27, 2023Updated 3 years ago
- Playing with CSM☆22Mar 14, 2025Updated last year
- run embeddings in MLX☆97Sep 27, 2024Updated last year
- jQuery, React and Streamlit applications written by LLMs☆16Dec 24, 2023Updated 2 years ago
- Paste Word, get Markdown☆17Jul 30, 2024Updated 2 years ago
- ☆21Oct 6, 2023Updated 2 years ago
- How to deploy PHP, nginx and Laravel queues via Coolify☆14Aug 30, 2025Updated 11 months ago