LLM 推理服务性能测试
☆45Dec 17, 2023Updated 2 years ago
Alternatives and similar repositories for llm-inference-benchmark
Users that are interested in llm-inference-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- survery of small language models☆18Jul 23, 2024Updated 2 years ago
- A Android client of Stable Diffusion.☆13Mar 29, 2024Updated 2 years ago
- ☆13Mar 29, 2026Updated 5 months ago
- Uniformaly: Towards Task-Agnostic Unified Anomaly Detection☆16Sep 15, 2023Updated 2 years ago
- AI拆解论文,人人都能读懂前沿研究☆24Jul 10, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- LLM Inference benchmark☆438Jul 23, 2024Updated 2 years ago
- High-Performance Linpack Benchmark adopted version for GPU backend☆12Sep 12, 2022Updated 3 years ago
- CUDA keyring packaging for Debian☆14Apr 14, 2023Updated 3 years ago
- An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.☆13Jun 7, 2023Updated 3 years ago
- Official implement of CIKM2025: 《UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion》☆21Sep 17, 2025Updated 11 months ago
- Build gstreamer on Raspberry Pi 3☆14Nov 2, 2018Updated 7 years ago
- MobileSAM のエンコーダー/デコーダーをONNXに変換し、推論するサンプル☆12Apr 11, 2024Updated 2 years ago
- ☆11Apr 10, 2023Updated 3 years ago
- N-body simulation based on CUDA.☆14Jun 20, 2019Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (N…☆12Jun 24, 2024Updated 2 years ago
- 「城语」APP基于A级景区、历史古迹、文物保护单位等基础数据,利用先进的大模型能力实现智能化的Citywalk 路线规划,包括设计一条路线、生成路线攻略、生成景点的推荐理由等三大核心功能;利用大模型减少了人工编辑和推荐的工作量,并可以根据游客的需求进行个性化定制,提升了游客…☆19Feb 20, 2024Updated 2 years ago
- Notes on putting micropython on STM32F407VG bare board☆11Oct 7, 2019Updated 6 years ago
- A simple, easy-to-hack GraphRAG implementation☆15Sep 21, 2024Updated last year
- Generate text images for training deep learning ocr model☆10Oct 22, 2018Updated 7 years ago
- LLM Agents: Landing Page Generation for an E-commerce Platform Using CrewAI, Groq-LangChain and Qdrant☆15May 30, 2024Updated 2 years ago
- ☆17Nov 27, 2023Updated 2 years ago
- 小飞机翻墙教程☆24Nov 14, 2019Updated 6 years ago
- ☆16Jan 23, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Visual self-questioning for large vision-language assistant.☆44Jul 23, 2025Updated last year
- Code for MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization☆23Feb 18, 2026Updated 6 months ago
- ☆25Mar 31, 2022Updated 4 years ago
- [ICML2026] FreeRet: MLLMs as Training-Free Retrievers☆25May 25, 2026Updated 3 months ago
- Pre-built ROCm-GDB and GPU Debug SDK binaries☆16Mar 21, 2019Updated 7 years ago
- Implementation of various algorithms in the Nested Sequential Monte Carlo family of methods.☆14Sep 9, 2015Updated 10 years ago
- 基于internlm-chat-7b的保险知识大模型微调☆20Apr 26, 2024Updated 2 years ago
- 该部分为自 己在学习tensorflow2.0中实现的各种模型还有算法,供大家参考☆19Jul 30, 2020Updated 6 years ago
- In this programming assignment you will implement a streaming video server and client that communicate control commands via the Real-Time…☆11Dec 29, 2012Updated 13 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- inference on tvm runtime using c++ with gpu enabled☆10Apr 25, 2018Updated 8 years ago
- ☆12Jan 25, 2023Updated 3 years ago
- ☆14Jun 11, 2024Updated 2 years ago
- Privacy-preserving k-means clustering on data owned by multiple parties☆14May 10, 2016Updated 10 years ago
- ☆17Jun 10, 2025Updated last year
- Systemback_source-1.9.4☆15Jan 2, 2021Updated 5 years ago
- 用大模型批量处理数据,现支持各种大模型做OCR,支持通义千问, 月之暗面, 百度飞桨OCR, OpenAI 和LLAVA。Use LLM to generate or clean data for academic use. Support OCR with qwen, m…☆18Sep 15, 2024Updated last year