a collection of skills for vllm-omni
☆88Sep 21, 2026Updated last week
Alternatives and similar repositories for vllm-omni-skills
Users that are interested in vllm-omni-skills are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Agent skills for vLLM☆103Apr 3, 2026Updated 5 months ago
- A framework for efficient model inference with omni-modality models☆7,084Updated this week
- vLLM Daily Summarization of Merged PRs☆54Updated this week
- SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.☆1,282Updated this week
- An LLM post-training framework with vLLM for RL Scaling☆475Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A lightweight `vLLM-Omni`-style diffusion implementation built around `Wan2.2-TI2V-5B-Diffusers` inspired from nano-vllm☆63May 25, 2026Updated 4 months ago
- Fast and memory-efficient exact attention☆23Jun 26, 2026Updated 3 months ago
- A unified VLA post-training framework for human-in-the-loop data collection, fine-tuning, and reinforcement learning.☆102Sep 11, 2026Updated 2 weeks ago
- TurboServe: Serving Streaming Video Generation Efficiently and Economically☆247Jul 12, 2026Updated 2 months ago
- ☆14Jan 27, 2026Updated 8 months ago
- A PyTorch Implementation of DF-GAN☆10Mar 26, 2022Updated 4 years ago
- ☆236Aug 26, 2026Updated last month
- ☆19Feb 18, 2025Updated last year
- A high-performance and light-weight router for vLLM large scale deployment☆438Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Voxtral Codec : Combining Semantic VQ and Acoustic FSQ for Ultra-Low Bitrate Speech Generation (Voxtral TTS Backbone)☆17Mar 27, 2026Updated 6 months ago
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆27Jul 11, 2026Updated 2 months ago
- DLBlas: clean and efficient kernels☆46Sep 17, 2026Updated last week
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆567Sep 8, 2026Updated 2 weeks ago
- Source code repository for ASPLOS '25 paper "Syno: Structured Synthesis for Neural Operators"☆15Aug 31, 2025Updated last year
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆28Updated this week
- detect face using MTCNN and tracking face with KCF/Kalman Filter. The assignment problem solved by Hungarian.☆12Nov 16, 2019Updated 6 years ago
- TokenSpeed is a speed-of-light LLM inference engine.☆2,180Updated this week
- the sharing platform of SOSD lab☆12Jun 8, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The l4t linux kernel☆10Dec 5, 2018Updated 7 years ago
- This is the original matlab version of MKCFup☆10Jan 23, 2019Updated 7 years ago
- ☆112Oct 16, 2025Updated 11 months ago
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- vLLM plugin for block-based diffusion language model (dLLM) support☆29May 25, 2026Updated 4 months ago
- ☆51Jul 27, 2026Updated 2 months ago
- [ACL 2024 Findings] Light-PEFT: Lightening Parameter-Efficient Fine-Tuning via Early Pruning☆13Sep 2, 2024Updated 2 years ago
- Machine Learning Phase of matter☆15Jun 12, 2019Updated 7 years ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,668Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated 2 months ago
- Omni inference in C/C++☆272Updated this week
- LightTTS is a lightweight TTS inference framework optimized for CosyVoice2 and CosyVoice3, enabling fast and scalable speech synthesis in…☆49Apr 14, 2026Updated 5 months ago
- [ICLR'25] ApolloMoE: Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts☆53Nov 20, 2024Updated last year
- Contains implementation of PINN using Tensorflow 2.4.0☆14Apr 28, 2023Updated 3 years ago
- This provider contains operators, decorators and triggers to send a ray job from an airflow task☆25Jul 1, 2026Updated 2 months ago
- ☆292Sep 5, 2026Updated 3 weeks ago