Summary of the Specs of Commonly Used GPUs for Training and Inference of LLM
☆79Aug 12, 2025Updated last year
Alternatives and similar repositories for GPUs-Specs
Users that are interested in GPUs-Specs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 9 months ago
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆25Jul 11, 2026Updated last month
- A simple calculation for LLM MFU.☆78Sep 10, 2025Updated 11 months ago
- Distributed Compiler and Optimized Parallel Kernels☆1,529Aug 12, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆82Updated this week
- ☆14Oct 3, 2024Updated last year
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- 🎓Automatically Update LLM inference systems Papers Daily using Github Actions (Update Every 12th hours)☆12Updated this week
- ☆15Feb 11, 2025Updated last year
- Cluster simulator with far memory☆12Apr 28, 2020Updated 6 years ago
- ☆12Dec 17, 2023Updated 2 years ago
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 3 months ago
- Compare different hardware platforms via the Roofline Model for LLM inference tasks.☆126Mar 13, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- PerFlow-AI is a programmable performance analysis, modeling, prediction tool for AI system.☆33May 12, 2026Updated 3 months ago
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,355Aug 28, 2025Updated last year
- repository for the MICCAI 2022 AutoPET challenge☆14Sep 19, 2022Updated 3 years ago
- ☆16Jul 28, 2021Updated 5 years ago
- A PyTorch wrapper of parallel exclusive scan in CUDA☆12May 25, 2023Updated 3 years ago
- ☆34Jul 13, 2026Updated last month
- Cycle-accurate C++ & SystemC simulator for the RISC-V GPGPU Ventus☆34May 17, 2026Updated 3 months ago
- High performance RDMA-based distributed feature collection component for training GNN model on EXTREMELY large graph☆55Jul 3, 2022Updated 4 years ago
- [ICML 2022] "Coarsening the Granularity: Towards Structurally Sparse Lottery Tickets" by Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wa…☆33Apr 9, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]☆67Mar 5, 2025Updated last year
- ⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.☆158May 10, 2025Updated last year
- Examples of CUDA implementations by Cutlass CuTe☆281Jul 1, 2025Updated last year
- Code for "Fast Sparse ConvNets" CVPR2020 submissions☆12Nov 20, 2019Updated 6 years ago
- DeepSeek-V3/R1 inference performance simulator☆195Mar 27, 2025Updated last year
- ☆101Apr 2, 2025Updated last year
- Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of pap…☆284Mar 6, 2025Updated last year
- Dynamic Memory Management for Serving LLMs without PagedAttention☆517Updated this week
- ☆37Aug 7, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,738Aug 13, 2026Updated 2 weeks ago
- Sequence-level 1F1B schedule for LLMs.☆36Aug 26, 2025Updated last year
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆550Updated this week
- ☆49Apr 15, 2024Updated 2 years ago
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆76Aug 8, 2026Updated 3 weeks ago
- Ventus GPGPU ISA Simulator Based on Spike☆53Aug 20, 2026Updated last week
- Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for…☆217Updated this week