Summary of the Specs of Commonly Used GPUs for Training and Inference of LLM
☆79Aug 12, 2025Updated last year
Alternatives and similar repositories for GPUs-Specs
Users that are interested in GPUs-Specs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 9 months ago
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆26Jul 11, 2026Updated 2 months ago
- A simple calculation for LLM MFU.☆78Sep 10, 2025Updated last year
- Distributed Compiler and Optimized Parallel Kernels☆1,546Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆14Oct 3, 2024Updated last year
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- 🎓Automatically Update LLM inference systems Papers Daily using Github Actions (Update Every 12th hours)☆12Updated this week
- ☆16Feb 11, 2025Updated last year
- Cluster simulator with far memory☆12Apr 28, 2020Updated 6 years ago
- ☆12Dec 17, 2023Updated 2 years ago
- ☆118Updated this week
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 3 months ago
- Compare different hardware platforms via the Roofline Model for LLM inference tasks.☆126Mar 13, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- PerFlow-AI is a programmable performance analysis, modeling, prediction tool for AI system.☆33May 12, 2026Updated 4 months ago
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,363Aug 28, 2025Updated last year
- repository for the MICCAI 2022 AutoPET challenge☆14Sep 19, 2022Updated 4 years ago
- ☆16Jul 28, 2021Updated 5 years ago
- A PyTorch wrapper of parallel exclusive scan in CUDA☆12May 25, 2023Updated 3 years ago
- ☆34Jul 13, 2026Updated 2 months ago
- Cycle-accurate C++ & SystemC simulator for the RISC-V GPGPU Ventus☆34May 17, 2026Updated 4 months ago
- High performance RDMA-based distributed feature collection component for training GNN model on EXTREMELY large graph☆55Jul 3, 2022Updated 4 years ago
- [ICML 2022] "Coarsening the Granularity: Towards Structurally Sparse Lottery Tickets" by Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wa…☆33Apr 9, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]☆68Mar 5, 2025Updated last year
- ⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.☆158May 10, 2025Updated last year
- Examples of CUDA implementations by Cutlass CuTe☆282Jul 1, 2025Updated last year
- Code for "Fast Sparse ConvNets" CVPR2020 submissions☆12Nov 20, 2019Updated 6 years ago
- DeepSeek-V3/R1 inference performance simulator☆195Mar 27, 2025Updated last year
- ☆101Apr 2, 2025Updated last year
- Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of pap…☆284Mar 6, 2025Updated last year
- Dynamic Memory Management for Serving LLMs without PagedAttention☆523Aug 24, 2026Updated 3 weeks ago
- ☆37Aug 7, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,783Aug 13, 2026Updated last month
- Sequence-level 1F1B schedule for LLMs.☆37Aug 26, 2025Updated last year
- ☆49Apr 15, 2024Updated 2 years ago
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆557Updated this week
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆77Aug 8, 2026Updated last month
- Ventus GPGPU ISA Simulator Based on Spike☆55Sep 3, 2026Updated 2 weeks ago
- ☆76Oct 31, 2024Updated last year