[ATC'25] Katz is a high-performance serving system designed specifically for diffusion model workflows with multiple adapters.
☆24May 26, 2025Updated last year
Alternatives and similar repositories for Katz
Users that are interested in Katz are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- End-to-end benchmark for AI-generated GPU kernels, drawn from real production traces — turn a PyTorch reference into a DSL kernel (Triton…☆28Updated this week
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆25Jul 11, 2026Updated last month
- Here are my personal paper reading notes (including machine learning systems, AI infrastructure, and other interesting stuffs).☆224Aug 6, 2026Updated 3 weeks ago
- Artifact for "Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving" [SOSP '24]☆24Nov 21, 2024Updated last year
- ☆12May 19, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 📚A curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.🎉☆590Jun 13, 2026Updated 2 months ago
- Error-free transformations are used to get results with extra accuracy.☆15May 27, 2026Updated 3 months ago
- COSE: Configuring Serverless Functions using Statistical Learning☆10Jun 28, 2023Updated 3 years ago
- tensorflow fork with Salus integration☆12Jan 7, 2022Updated 4 years ago
- A Distributed Analysis and Benchmarking Framework for Apache OpenWhisk Serverless Platform☆12Dec 11, 2018Updated 7 years ago
- PyTorch implementation of PTQ4DiT https://arxiv.org/abs/2405.16005☆49Nov 8, 2024Updated last year
- [ICML 2026] Official repository for the paper "Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention"☆49Aug 5, 2026Updated 3 weeks ago
- Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism (NIPS'25)☆16Oct 6, 2025Updated 10 months ago
- Official Code For Dual Grained Quantization: Efficient Fine-Grained Quantization for LLM☆14Dec 27, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Kratos: An FPGA Benchmark for Unrolled Deep Neural Networks with Fine-Grained Sparsity and Mixed Precision☆12Jan 19, 2026Updated 7 months ago
- NS3 Network routing simulation. SJTU CS339 Project☆11Dec 7, 2018Updated 7 years ago
- [ASPLOS' 26] TetriServe: Efficiently Serving Mixed DiT Workloads☆18Mar 12, 2026Updated 5 months ago
- Open-source implementation for "Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow"☆94Oct 15, 2025Updated 10 months ago
- The code for our paper "Neural Architecture Search as Program Transformation Exploration"