Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parallel architecture, suitable for MOE model hybrid inference.
☆111Aug 17, 2026Updated this week
Alternatives and similar repositories for Lsglang
Users that are interested in Lsglang are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆444Updated this week
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded …☆90Updated this week
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆23Jul 2, 2026Updated last month
- forked from vllm-project/flash-attention☆62May 9, 2026Updated 3 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A source repo of Postgres Chinese full-test search docker image, based on zhparser.☆10Mar 25, 2021Updated 5 years ago
- ☆11Updated this week
- RIFT 见缝插帧☆17May 15, 2022Updated 4 years ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 5 months ago
- a demo of Animeganv2 infer by ncnn☆23Nov 26, 2021Updated 4 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,077Aug 15, 2026Updated last week
- 检测透视图像中的矩 形文档并对其进行矫正☆32Sep 16, 2022Updated 3 years ago
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- Use ‘DICE-Talk’ in ComfyUI,which is a method about 'Correlation-Aware Emotional Talking Portrait Generation'.☆24May 7, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V…☆660Updated this week
- ☆25Jun 16, 2026Updated 2 months ago
- Hopenet: deep head pose estimator on ncnn☆10Jun 18, 2020Updated 6 years ago
- 在小爱音箱上获得与豆包近乎一致的端侧实时语音对话体验☆26Feb 26, 2026Updated 5 months ago
- ☆17Jun 13, 2026Updated 2 months ago
- ☆11Aug 11, 2026Updated last week
- 基于传统方法(非深度学习)的美颜算法实现,支持美白和磨皮☆13Jul 31, 2023Updated 3 years ago
- Triton for AMD MI25/50/60. Development repository for the Triton language and compiler☆35Dec 15, 2025Updated 8 months ago
- Please refer to the develop branch in https://github.com/jvcleave/ofxImGui. I'll keep this fork in sync until it's merged in master.☆15Apr 23, 2026Updated 4 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- qrencode for openresty☆11Sep 21, 2017Updated 8 years ago
- memo☆13Dec 22, 2022Updated 3 years ago
- deepin community SIG and repo management☆22Aug 12, 2026Updated last week
- Analysis of 44 AI agent frameworks through a context engineering lens. Feb 2026.☆17Feb 18, 2026Updated 6 months ago
- Building a paint app with Stable Diffusion running locally☆12Dec 4, 2022Updated 3 years ago
- for calculating euler angle test in ncnn☆11Aug 7, 2026Updated 2 weeks ago
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- ☆10Feb 18, 2024Updated 2 years ago
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-reques…☆624Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Lua XPath resurrection☆14Jul 12, 2024Updated 2 years ago
- A Lua XPath library based on libxml2.☆10Sep 16, 2016Updated 9 years ago
- 这个工具基于小智AI的MCP功能进行开发,允许通过 MCP 协议创建和查询苹果提醒事项(苹果备忘录),可以设置提醒内容(标题和备注)以及提醒日期时间。☆15May 26, 2025Updated last year
- Tensor parallelism is all you need. Run LLMs on an AI cluster at home using any device. Distribute the workload, divide RAM usage, and in…☆18Nov 11, 2024Updated last year
- Consul Events HTTP API Wrapper☆14Sep 15, 2021Updated 4 years ago
- openresty luajit ffi bindings for libbzip2 - bzip2 compress library☆12Nov 10, 2016Updated 9 years ago
- Lua bitmaps (aka bitstrings or bitsets) implemented in C☆15Mar 6, 2014Updated 12 years ago