a collection of skills for vllm-omni
☆86Sep 7, 2026Updated this week
Alternatives and similar repositories for vllm-omni-skills
Users that are interested in vllm-omni-skills are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Agent skills for vLLM☆96Apr 3, 2026Updated 5 months ago
- A framework for efficient model inference with omni-modality models☆6,702Updated this week
- vLLM Daily Summarization of Merged PRs☆54Updated this week
- SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.☆1,106Updated this week
- An LLM post-training framework with vLLM for RL Scaling☆454Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- TurboServe: Serving Streaming Video Generation Efficiently and Economically☆43Jul 12, 2026Updated last month
- A lightweight `vLLM-Omni`-style diffusion implementation built around `Wan2.2-TI2V-5B-Diffusers` inspired from nano-vllm☆61May 25, 2026Updated 3 months ago
- ☆164Mar 5, 2026Updated 6 months ago
- Fast and memory-efficient exact attention☆23Jun 26, 2026Updated 2 months ago
- A unified VLA post-training framework for human-in-the-loop data collection, fine-tuning, and reinforcement learning.☆79Sep 1, 2026Updated last week
- ☆14Jan 27, 2026Updated 7 months ago
- On-device speech AI runtime for ASR, TTS, VAD, and voice cloning. Python-simple, C++-native, GGUF-powered.☆27Aug 7, 2026Updated last month
- A simple tomasulo simulator written in Rust for the course Computer Architecture.☆14Dec 29, 2022Updated 3 years ago
- PyTorch Tutorial to train ConvNets for Image Classification.☆11May 20, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆220Aug 26, 2026Updated last week
- ☆19Feb 18, 2025Updated last year
- A high-performance and light-weight router for vLLM large scale deployment☆396Aug 31, 2026Updated last week
- Voxtral Codec : Combining Semantic VQ and Acoustic FSQ for Ultra-Low Bitrate Speech Generation (Voxtral TTS Backbone)☆17Mar 27, 2026Updated 5 months ago
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆25Jul 11, 2026Updated last month
- DLBlas: clean and efficient kernels☆47Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆538Jul 15, 2026Updated last month
- Source code repository for ASPLOS '25 paper "Syno: Structured Synthesis for Neural Operators"☆15Aug 31, 2025Updated last year
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆28Apr 8, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆29Jul 4, 2026Updated 2 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆2,103Updated this week
- Finetuning LLaMA with DeepSpeed☆10Apr 14, 2023Updated 3 years ago
- ☆110Oct 16, 2025Updated 10 months ago
- FastLongSpeech is a novel framework designed to extend the capabilities of Large Speech-Language Models for efficient long-speech process…☆16Jul 22, 2025Updated last year
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- Dataset for Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models in Interspeech 2024.☆16Jul 4, 2024Updated 2 years ago
- vLLM plugin for block-based diffusion language model (dLLM) support☆29May 25, 2026Updated 3 months ago
- [ACL 2024 Findings] Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruning☆13Sep 2, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A better wrapper for using RDMA programming APIs in Rust flavor☆95Aug 13, 2026Updated 3 weeks ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,516Updated this week
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated last month
- Code for the paper "Knowledge-Aware Federated Active Learning with Non-IID Data", ICCV2023☆10Sep 8, 2023Updated 3 years ago
- ☆270Updated this week
- [ICLR'25] ApolloMoE: Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts☆53Nov 20, 2024Updated last year
- source code of EfficientTTS 2☆22Feb 18, 2024Updated 2 years ago