The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"
☆18Mar 17, 2025Updated last year
Alternatives and similar repositories for dynamic-batching
Users that are interested in dynamic-batching are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Burstable Cloud Scheduler☆17Jun 6, 2024Updated 2 years ago
- [NAACL 2025 Main Selected Oral] Repository for the paper: Prompt Compression for Large Language Models: A Survey☆36May 18, 2025Updated last year
- ☆17May 10, 2024Updated 2 years ago
- Computational Memory Neural Network Compiler☆11Aug 11, 2021Updated 4 years ago
- A variant of Ahash written in C++.☆10Mar 20, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- Benchmarking Machine Learning Model Inference in Data Streaming Solutions☆10Jun 12, 2024Updated 2 years ago
- MNSIM version 1.1. We have uploaded a high-level modeling tool and please use this version: https://github.com/Zhu-Zhenhua/MNSIM_Python☆12Dec 12, 2019Updated 6 years ago
- Benchmark framework of compute-in-memory based accelerators for deep neural network (inference engine focused)☆11Jun 1, 2021Updated 5 years ago
- Codes for our paper "Exploring Bit-Slice Sparsity in Deep Neural Networks for Efficient ReRAM-Based Deployment" [NeurIPS'19 EMC2 workshop]…☆10Oct 12, 2020Updated 5 years ago
- DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting☆19Mar 4, 2025Updated last year
- Java-like Language with Static Information Flow Types☆14May 5, 2025Updated last year
- ☆22Jun 6, 2022Updated 4 years ago
- 컴퓨터 신기술 특강☆13Dec 22, 2018Updated 7 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A benchmark framework for decision forest inferences☆12Jan 19, 2024Updated 2 years ago
- Apache DataFusion Benchmarks☆23May 2, 2026Updated 3 months ago
- Neural Network-Hardware Co-design for Scalable RRAM-based BNN Accelerators☆11Apr 9, 2019Updated 7 years ago
- ☆32Jan 16, 2025Updated last year
- ☆10Sep 26, 2024Updated last year
- Artifact for "Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs" [arXiv '25]☆21Jul 26, 2026Updated 2 weeks ago
- Continuous Pipelined Speculative Decoding☆22May 25, 2026Updated 2 months ago
- ☆11Jan 12, 2021Updated 5 years ago
- An LLM inference engine, written in C++☆20Mar 30, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The speculative decoding for uav environment☆17Oct 18, 2025Updated 9 months ago
- Code for our paper "Evaluating SIMD Compiler-Intrinsics for Database Systems"☆16Jul 5, 2023Updated 3 years ago
- A novel system that unifies LLM serving with query optimization to efficiently process batch agentic workflows.☆16Jun 14, 2026Updated last month
- SEU网络自动重连☆13May 24, 2019Updated 7 years ago
- Benchmarking Semantic Query Processing Engines☆63Jul 16, 2026Updated 3 weeks ago
- fat-tree topology and routing algorithms using mininet and pox☆14Jan 30, 2019Updated 7 years ago
- ☆16Jan 14, 2025Updated last year
- Make reasoning models scalable☆51Jun 2, 2026Updated 2 months ago
- This Dataset consists of heterogeneous IoT jobs/tuples containing the information - TupleName, TupleId, Size of the Tuple (Size of a job)…☆16Oct 28, 2019Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Bayesian Wi-Fi rate control☆21Nov 30, 2016Updated 9 years ago
- Official Implementation of Paper FOLDER (ICCV2025) and Turbo (ECCV2024)☆15Jun 27, 2025Updated last year
- Architecture for RRAM multilevel programming☆18Sep 6, 2018Updated 7 years ago
- ☆13Jun 15, 2020Updated 6 years ago
- 一个用于练习 BMAD 开发工作流的学习项目☆17Apr 14, 2026Updated 3 months ago
- ☆18Jun 17, 2022Updated 4 years ago
- ☆12May 23, 2022Updated 4 years ago