Local LLM Inference Speed Test Tool
☆145Jul 30, 2026Updated last week
Alternatives and similar repositories for llm_speedtest
Users that are interested in llm_speedtest are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V…☆586Updated this week
- forked from vllm-project/flash-attention☆59May 9, 2026Updated 3 months ago
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode…☆483Jul 13, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 大语言模型工具集☆28Aug 1, 2025Updated last year
- TensorRT depth-anything for anyone and anywhere☆16Jan 29, 2024Updated 2 years ago
- Run DeepSeek-V4-Flash on SM89 (Ada / RTX 4090) with vLLM — patch over PR #41834. Validated on 4x RTX 4090.☆95Updated this week
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆21Jul 2, 2026Updated last month
- 本项目用于将 Docker 镜像同步到 GitHub 的仓库 ghcr.io 以加速国内 Docker 镜像的下载,若ghcr.io下载速度慢可尝试使用南京大学加速站 ghcr.nju.edu.cn 加速下载。☆16Updated this week
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆30Jan 22, 2026Updated 6 months ago
- A fork of vLLM enabling Pascal architecture GPUs☆35Feb 21, 2025Updated last year
- ⚡️Qwen-Image 4.8x🎉 speedup with Hybrid Acceleration for low VRAM GPUs☆17Oct 24, 2025Updated 9 months ago
- 华为软件大赛: 最小费用流(zkw)+遗传算法。 zkw比sfa的费用流快好多啊,10倍左右,牛逼~☆13Apr 6, 2017Updated 9 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- paper-read-notes☆13Sep 26, 2024Updated last year
- A lightweight, production-ready C++ library for LLM tokenization, fully compatible with HuggingFace tokenizer.json.☆33Jan 4, 2026Updated 7 months ago
- LLM Inference FrameWork☆32May 20, 2026Updated 2 months ago
- ☆32Aug 3, 2026Updated last week
- ncnn 实现一些项目例子☆27Feb 17, 2023Updated 3 years ago
- learn TensorRT from scratch🥰☆18Sep 29, 2024Updated last year
- 搜藏的希望的代码片段☆13Jun 6, 2023Updated 3 years ago
- HunyuanDiT with TensorRT and libtorch☆18May 22, 2024Updated 2 years ago
- a ai infra framework for edge device base on nndeploy☆18Nov 27, 2025Updated 8 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- 使用mnn-llm对GOT-OCR2.0进行推理☆13Oct 2, 2024Updated last year
- Implementation of a histogram equalization program using CUDA. Histogram equalization is a technique for adjusting image intensities to e…☆13Jan 3, 2021Updated 5 years ago
- Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free…☆24Updated this week
- This project is primarily used to deploy large language models and multimodal large models on Orin.🚀🚀🚀☆18Jun 23, 2026Updated last month
- The LLM API Benchmark Tool is a flexible Go-based utility designed to measure and analyze the performance of OpenAI-compatible API endpoi…☆96Mar 25, 2026Updated 4 months ago
- ☆17Nov 14, 2023Updated 2 years ago
- 跟着Tensorrt_pro学习各种知识☆39Nov 25, 2022Updated 3 years ago
- segmentation algorithm yolact use tensorrt deploy☆14May 7, 2022Updated 4 years ago
- Burrows-Wheeler Aligner for x86,x86_64, arm and aarch64 architectures (PC, Raspberry PI, ODROID, M1)☆10Apr 13, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A lightweight, single-header C++11 Jinja2 template engine for LLM chat templates.☆20Mar 4, 2026Updated 5 months ago
- Inference Llama 2 in one file of pure Cuda☆17Aug 20, 2023Updated 2 years ago
- Detection and Tracking ROS node based on CenterPoint and Kalman Filter☆24Feb 24, 2024Updated 2 years ago
- ☆60Dec 9, 2025Updated 8 months ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated 11 months ago
- 基于 esp8266 的智能电表功率计☆12Apr 28, 2022Updated 4 years ago
- A curated collection of reusable AI Agent Skills for standardized workflows, best practices, and domain expertise.☆22May 29, 2026Updated 2 months ago