Local LLM Inference Speed Test Tool
☆226Sep 2, 2026Updated 2 weeks ago
Alternatives and similar repositories for llm_speedtest
Users that are interested in llm_speedtest are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,102Updated this week
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆457Updated this week
- forked from vllm-project/flash-attention☆65May 9, 2026Updated 4 months ago
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 200+ tok/s single-reques…☆906Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- TensorRT depth-anything for anyone and anywhere☆16Jan 29, 2024Updated 2 years ago
- Run DeepSeek-V4.1-Flash / DeepSeek-V4-Flash and GLM-5.3-Flash on SM89 (Ada / RTX 4090) and SM120 (RTX PRO 6000) with vLLM☆164Updated this week
- ⚡️Qwen-Image 4.8x🎉 speedup with Hybrid Acceleration for low VRAM GPUs☆17Oct 24, 2025Updated 10 months ago
- Flash Attention in ~100 lines of CUDA (forward pass only)☆12Jun 10, 2024Updated 2 years ago
- paper-read-notes☆13Sep 26, 2024Updated last year
- Unlock P2P comms between consumer NVIDIA GPUs☆539Updated this week
- ☆16Mar 9, 2026Updated 6 months ago
- 搜藏的希望的代码片段☆13Jun 6, 2023Updated 3 years ago
- HunyuanDiT with TensorRT and libtorch☆18May 22, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- a ai infra framework for edge device base on nndeploy☆18Nov 27, 2025Updated 9 months ago
- 使用mnn-llm对GOT-OCR2.0进行推理☆13Oct 2, 2024Updated last year
- Implementation of a histogram equalization program using CUDA. Histogram equalization is a technique for adjusting image intensities to e…☆13Jan 3, 2021Updated 5 years ago
- 加拿大的社会新闻生活信息☆11Mar 24, 2023Updated 3 years ago
- ☆17Nov 14, 2023Updated 2 years ago
- Burrows-Wheeler Aligner for x86,x86_64, arm and aarch64 architectures (PC, Raspberry PI, ODROID, M1)☆10Apr 13, 2022Updated 4 years ago
- A lightweight, single-header C++11 Jinja2 template engine for LLM chat templates.☆20Mar 4, 2026Updated 6 months ago
- Ontology Modeling Language (OML) Workbench☆14Mar 12, 2020Updated 6 years ago
- Inference Llama 2 in one file of pure Cuda☆17Aug 20, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Detection and Tracking ROS node based on CenterPoint and Kalman Filter☆24Feb 24, 2024Updated 2 years ago
- iptables-trace is an eBPF enhanced iptables-TRACE alternative iptables TRACE. GPL-3.0 license☆14Feb 3, 2025Updated last year
- ☆61Dec 9, 2025Updated 9 months ago
- A curated collection of reusable AI Agent Skills for standardized workflows, best practices, and domain expertise.☆22May 29, 2026Updated 3 months ago
- A ESP32 BLE scanner with iotWebConf and MQTT☆13Mar 30, 2019Updated 7 years ago
- ☆53Apr 16, 2026Updated 5 months ago
- 基于 C++23 的模块化 LLM 推理框架,原生解析 GGUF 格式。☆48Jun 19, 2026Updated 3 months ago
- fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tp…☆5,066Updated this week
- ☆35Jul 2, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A one-page-only CGraph-API-liked DAG project.☆28Feb 11, 2025Updated last year
- People and asset tracking system. Hardware support : 1. Happy Bubble Presence Detectors 2.Raspberryp Pi 3. ESP32 kit with Wifi and Blueto…☆10Jun 6, 2018Updated 8 years ago
- A Minimalistic Auto-Diff Optimization Framework for Teaching and Understanding Pytorch☆28Updated this week
- ☆71Mar 8, 2026Updated 6 months ago
- LuatOS flash tool for linux -- utilities to flash soc file, generate and flash script.img of LuatOS to OpenLuat air101/air103 and esp32s…☆13Mar 16, 2023Updated 3 years ago
- ☆26Aug 15, 2023Updated 3 years ago
- ☆32Aug 25, 2023Updated 3 years ago