Dynamic batching library for Deep Learning inference. Tutorials for LLM, GPT scenarios.
☆106Aug 24, 2026Updated last month
Alternatives and similar repositories for batch-inference
Users that are interested in batch-inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Predict the performance of LLM inference services☆23Sep 18, 2025Updated last year
- aigc evals☆10Dec 2, 2023Updated 2 years ago
- pytorch版bert权重转tf☆22May 19, 2020Updated 6 years ago
- MetaMask like widget for Binance chain☆15Sep 8, 2022Updated 4 years ago
- Generate tweets with Titter GPT, an easy to use streamlit app leveraging OpenAI's GPT-3 model and revGPT. Create custom tweet styles, ton…☆24May 16, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A time delay estimation method for event-based time-series data. Time delay estimation is also known as the correction of time offsets an…☆16Dec 3, 2025Updated 10 months ago
- ☆18Dec 27, 2023Updated 2 years ago
- LLM Serving Performance Evaluation Harness☆86Feb 25, 2025Updated last year
- ☆18Jan 2, 2024Updated 2 years ago
- ☆11Jul 3, 2023Updated 3 years ago
- 🗣️ Convert between phonetic alphabets☆11Feb 7, 2022Updated 4 years ago
- Python Inference Script(PyIS)☆19Aug 30, 2022Updated 4 years ago
- Simple and easy stable diffusion inference with LightningModule on GPU, CPU and MPS (Possibly all devices supported by Lightning).☆16Jul 27, 2023Updated 3 years ago
- Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion☆20Jul 9, 2019Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.☆2,114Jun 30, 2025Updated last year
- accelerate generating vector by using onnx model☆18Jan 23, 2024Updated 2 years ago
- ☆13Jun 20, 2025Updated last year
- Benchmark for machine learning model online serving (LLM, embedding, Stable-Diffusion, Whisper)☆28Jun 28, 2023Updated 3 years ago
- ☆32Jan 30, 2023Updated 3 years ago
- [INTERSPEECH 2026] Official code for "Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech"☆21Sep 20, 2026Updated 2 weeks ago
- golang vad (voice activity detection) library based on webrtc☆13Dec 13, 2021Updated 4 years ago
- ☆18Oct 31, 2022Updated 3 years ago
- ☆22Dec 3, 2021Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Entity Linking within a Social Media Platform☆11May 2, 2019Updated 7 years ago
- Summary of system papers/frameworks/codes/tools on training or serving large model☆57Dec 17, 2023Updated 2 years ago
- ☆17Jun 30, 2020Updated 6 years ago
- ☆126Jun 25, 2026Updated 3 months ago
- ☆20May 13, 2022Updated 4 years ago
- StrategyQA 데이터 세트 번역☆22Apr 12, 2024Updated 2 years ago
- ☆16Nov 19, 2023Updated 2 years ago
- experiments with inference on llama☆103Jun 6, 2024Updated 2 years ago
- Just simple JavaScript framework. Provides support for manipulating with DOM and events handling. Easy for use, optimized for performance…☆11Feb 15, 2017Updated 9 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository showcases how to implement trunk-based development workflow while working in a Machine Learning project.☆42Sep 16, 2022Updated 4 years ago
- Common source, scripts and utilities for creating Triton backends.☆380Oct 2, 2026Updated last week
- High-performance vector search engine with no loss of accuracy through GPU and dynamic placement☆33Jul 12, 2025Updated last year
- Code for the paper "RIR-in-a-Box : Estimating Room Acoustics from 3D Mesh Data through Shoebox Approximation" presented at Interspeech 20…☆16Sep 1, 2024Updated 2 years ago
- spotify cli for the official client via dbus☆13May 5, 2020Updated 6 years ago
- Progressively growing of GANs Pytorch Implementation☆13Nov 9, 2017Updated 8 years ago
- Sekai Viewer but built with Next, optimized for performance☆11Jan 20, 2023Updated 3 years ago