a collection of skills for vllm-omni
☆84Aug 10, 2026Updated last week
Alternatives and similar repositories for vllm-omni-skills
Users that are interested in vllm-omni-skills are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A framework for efficient model inference with omni-modality models☆6,150Updated this week
- vLLM Daily Summarization of Merged PRs☆53Updated this week
- SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.☆835Updated this week
- TurboServe: Serving Streaming Video Generation Efficiently and Economically☆42Jul 12, 2026Updated last month
- Easily deploy your rwkv model☆19May 5, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A lightweight `vLLM-Omni`-style diffusion implementation built around `Wan2.2-TI2V-5B-Diffusers` inspired from nano-vllm☆58May 25, 2026Updated 2 months ago
- ☆162Mar 5, 2026Updated 5 months ago
- A unified VLA post-training framework for human-in-the-loop data collection, fine-tuning, and reinforcement learning.☆54Updated this week
- Fast and memory-efficient exact attention☆23Jun 26, 2026Updated last month
- A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.☆1,250Updated this week
- On-device speech AI runtime for ASR, TTS, VAD, and voice cloning. Python-simple, C++-native, GGUF-powered.☆25Aug 7, 2026Updated last week
- ☆187May 24, 2026Updated 2 months ago
- Stateful API logic for agentic applications using vLLM☆70Updated this week
- ☆19Feb 18, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A high-performance and light-weight router for vLLM large scale deployment☆362Updated this week
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆25Jul 11, 2026Updated last month
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆515Jul 15, 2026Updated last month
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆26Apr 8, 2026Updated 4 months ago
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆27Jul 4, 2026Updated last month
- TokenSpeed is a speed-of-light LLM inference engine.☆1,926Updated this week
- Finetuning LLaMA with DeepSpeed☆10Apr 14, 2023Updated 3 years ago
- FastLongSpeech is a novel framework designed to extend the capabilities of Large Speech-Language Models for efficient long-speech process…☆16Jul 22, 2025Updated last year
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆13Apr 17, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Dataset for Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models in Interspeech 2024.☆16Jul 4, 2024Updated 2 years ago
- vLLM plugin for block-based diffusion language model (dLLM) support☆27May 25, 2026Updated 2 months ago
- ☆48Jul 27, 2026Updated 3 weeks ago
- [ACL 2024 Findings] Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruning☆13Sep 2, 2024Updated last year
- An Efficient Supply Chain Management System using Blockchain & Machine Learning.☆10Nov 27, 2019Updated 6 years ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,313Updated this week
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆132Jul 25, 2026Updated 3 weeks ago
- Code for the paper "Knowledge-Aware Federated Active Learning with Non-IID Data", ICCV2023☆10Sep 8, 2023Updated 2 years ago
- source code of EfficientTTS 2☆21Feb 18, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Generation of Debian rootfs for multiple architectures☆14Nov 13, 2021Updated 4 years ago
- An ITK implementation of the GraphCut framework. See 'Graph cuts and efficient ND image segmentation' by Boykov and Funka-Lea and 'Intera…☆12Sep 18, 2017Updated 8 years ago
- verl Zero-Mismatch Dense/MoE HuggingFace Rollout☆64Aug 5, 2026Updated 2 weeks ago
- Rhetorical sentence classification using LLMs☆11Oct 26, 2025Updated 9 months ago
- Offical implementation of "MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map" (NeurIPS2024 Oral)☆36Jan 18, 2025Updated last year
- ☆34Jul 27, 2026Updated 3 weeks ago
- This project leverages advanced AI agents from crewAI to assist doctors in diagnosing medical conditions and recommending treatment plans…☆15Nov 16, 2024Updated last year