vLLM plugin for block-based diffusion language model (dLLM) support
☆28May 25, 2026Updated 4 months ago
Alternatives and similar repositories for dllm-plugin
Users that are interested in dllm-plugin are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- How to plot for papers, slides, demos, etc.☆10Apr 7, 2022Updated 4 years ago
- vLLM Daily Summarization of Merged PRs☆54Updated this week
- [ICLR 2026] AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size☆17Jan 28, 2026Updated 8 months ago
- Claude Code skill: Generate file-by-file code tutorial websites for any repository with parallel agent teams☆28Mar 13, 2026Updated 6 months ago
- High-performance Rust benchmark client for vLLM serving endpoints.☆51Aug 3, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 내맘대로 alluxio 정리중☆11May 13, 2019Updated 7 years ago
- ☆37Sep 13, 2025Updated last year
- 通过实验对比LLM推理中Prefill和Decoding阶段的吞吐量差异,揭示性能瓶颈,解释PD分离优化技术的原理。包含CUDA和Apple MPS (M系列芯片) 的测试脚本。☆23May 22, 2025Updated last year
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- A repository for Mixed Precision Training in JAX based on Equinox.☆16Nov 2, 2025Updated 11 months ago
- SUSTech CSE Courses☆11May 31, 2021Updated 5 years ago
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models (ICLR2026)☆23Jun 24, 2026Updated 3 months ago
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆28Jul 11, 2026Updated 2 months ago
- Evaluate state-of-the-art GPU joins☆14Nov 29, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- An LLM post-training framework with vLLM for RL Scaling☆480Updated this week
- Implementation of ICCV 2025 paper "Growing a Twig to Accelerate Large Vision-Language Models".☆32May 23, 2026Updated 4 months ago
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆105Apr 7, 2026Updated 5 months ago
- Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources☆56Sep 1, 2026Updated last month
- Training an LSTM network on the Penn Tree Bank (PTB) dataset☆11Nov 5, 2018Updated 7 years ago
- ☆23May 5, 2026Updated 4 months ago
- This is a fork of SGLang for hip-attention integration. Please refer to hip-attention for detail.☆18Mar 31, 2026Updated 6 months ago
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 4 months ago
- ☆20Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- MLIR-based TileLang Ascend Adapter☆27Sep 24, 2026Updated last week
- A code-generating database system with incorporated versioning commands in SQL.☆13Jan 18, 2021Updated 5 years ago
- Segmented Code Adjustment Quantization (SAQ)☆28Sep 22, 2025Updated last year
- TurboServe: Serving Streaming Video Generation Efficiently and Economically☆246Jul 12, 2026Updated 2 months ago
- ☆34Updated this week
- A Rust style C++ library.☆19Sep 3, 2022Updated 4 years ago
- Artifacts of EuroSys'24 paper "Exploring Performance and Cost Optimization with ASIC-Based CXL Memory"☆31Feb 21, 2024Updated 2 years ago
- [ICLR2025] Breaking Throughput-Latency Trade-off for Long Sequences with Speculative Decoding☆159Dec 4, 2024Updated last year
- ☆14Mar 7, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Sep 1, 2021Updated 5 years ago
- A code base for Vexless☆17Mar 7, 2024Updated 2 years ago
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆55Oct 18, 2024Updated last year
- ☆15Dec 2, 2019Updated 6 years ago
- ☆28Jun 8, 2026Updated 3 months ago
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- Memory-Bounded GPU Acceleration for Vector Search☆33Dec 29, 2025Updated 9 months ago