vLLM plugin for block-based diffusion language model (dLLM) support
☆27May 25, 2026Updated 2 months ago
Alternatives and similar repositories for dllm-plugin
Users that are interested in dllm-plugin are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- How to plot for papers, slides, demos, etc.☆10Apr 7, 2022Updated 4 years ago
- vLLM plugin for HyperCLOVAX☆15Jan 27, 2026Updated 6 months ago
- Agent skills for vLLM☆91Apr 3, 2026Updated 4 months ago
- vLLM Daily Summarization of Merged PRs☆52Updated this week
- [ICLR 2026] AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size☆15Jan 28, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- High-performance Rust benchmark client for vLLM serving endpoints.☆51Updated this week
- Easy, Fast, and Scalable Multimodal AI☆130Jun 2, 2026Updated 2 months ago
- A Python tool for tracking changes in Compute Express Link (CXL) features within the Linux kernel using GitHub API. It supports various o…☆13Jun 23, 2026Updated last month
- ☆34Sep 13, 2025Updated 10 months ago
- 通过实验对比LLM推理中Prefill和Decoding阶段的吞吐量差异,揭示性能瓶颈,解释PD分离优化技术的原理。包含CUDA和Apple MPS (M系列芯片) 的测试脚本。☆23May 22, 2025Updated last year
- StartupHeroes Checkstyle project with additional Checkstyle checks and Sonar Checkstyle plugin☆10Jan 25, 2024Updated 2 years ago
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- Examples of App of Apps Pattern☆10Jan 17, 2023Updated 3 years ago
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models (ICLR2026)☆23Jun 24, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Content for Tableau's OSS projects index☆11Jul 2, 2026Updated last month
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆25Jul 11, 2026Updated 3 weeks ago
- Evaluate state-of-the-art GPU joins☆14Nov 29, 2023Updated 2 years ago
- TurboServe: Serving Streaming Video Generation Efficiently and Economically☆38Jul 12, 2026Updated 3 weeks ago
- An LLM post-training framework with vLLM for RL Scaling☆400Updated this week
- An optimized Merkle Patricia Trie implementation on GPU, fully compatible with and integrable into Ethereum. The paper is published on VL…☆14Apr 15, 2024Updated 2 years ago
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆106Apr 7, 2026Updated 3 months ago
- Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources☆53Jul 19, 2026Updated 2 weeks ago
- Training an LSTM network on the Penn Tree Bank (PTB) dataset☆11Nov 5, 2018Updated 7 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ODQA Baseline 팀프로젝트 이슈/정보 저장용 레포입니다.☆12May 22, 2021Updated 5 years ago
- ☆23May 5, 2026Updated 2 months ago
- This is a fork of SGLang for hip-attention integration. Please refer to hip-attention for detail.☆18Mar 31, 2026Updated 4 months ago
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 2 months ago
- MLIR-based TileLang Ascend Adapter☆22Jul 25, 2026Updated last week
- Segmented Code Adjustment Quantization (SAQ)☆26Sep 22, 2025Updated 10 months ago
- A Rust style C++ library.☆19Sep 3, 2022Updated 3 years ago
- [ICLR2025] Breaking Throughput-Latency Trade-off for Long Sequences with Speculative Decoding☆156Dec 4, 2024Updated last year
- ☆14Mar 7, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- dInfer: An Efficient Inference Framework for Diffusion Language Models☆476Feb 11, 2026Updated 5 months ago
- ☆15Mar 3, 2024Updated 2 years ago
- ☆12Sep 1, 2021Updated 4 years ago
- kubeflow example☆18Jun 26, 2021Updated 5 years ago
- Helm chart for Flask App☆16Nov 6, 2022Updated 3 years ago
- Accelerating GPU Data Processing using FastLanes Compression☆19May 9, 2024Updated 2 years ago
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆54Oct 18, 2024Updated last year