Summary of the Specs of Commonly Used GPUs for Training and Inference of LLM
☆79Aug 12, 2025Updated 11 months ago
Alternatives and similar repositories for GPUs-Specs
Users that are interested in GPUs-Specs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated 11 months ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 8 months ago
- [ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation☆25Jul 11, 2026Updated 3 weeks ago
- A simple calculation for LLM MFU.☆78Sep 10, 2025Updated 11 months ago
- Distributed Compiler based on Triton for Parallel Systems☆1,512Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆60Updated this week
- ☆14Oct 3, 2024Updated last year
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- 🎓Automatically Update LLM inference systems Papers Daily using Github Actions (Update Every 12th hours)☆12Aug 3, 2026Updated last week
- ☆15Feb 11, 2025Updated last year
- Cluster simulator with far memory☆12Apr 28, 2020Updated 6 years ago
- ☆12Dec 17, 2023Updated 2 years ago
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 2 months ago
- Compare different hardware platforms via the Roofline Model for LLM inference tasks.☆125Mar 13, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- PerFlow-AI is a programmable performance analysis, modeling, prediction tool for AI system.☆33May 12, 2026Updated 2 months ago
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,352Aug 28, 2025Updated 11 months ago
- repository for the MICCAI 2022 AutoPET challenge☆14Sep 19, 2022Updated 3 years ago
- ☆16Jul 28, 2021Updated 5 years ago
- A PyTorch wrapper of parallel exclusive scan in CUDA☆12May 25, 2023Updated 3 years ago
- ☆33Jul 13, 2026Updated 3 weeks ago
- Cycle-accurate C++ & SystemC simulator for the RISC-V GPGPU Ventus☆34May 17, 2026Updated 2 months ago
- High performance RDMA-based distributed feature collection component for training GNN model on EXTREMELY large graph☆55Jul 3, 2022Updated 4 years ago
- [ICML 2022] "Coarsening the Granularity: Towards Structurally Sparse Lottery Tickets" by Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wa…☆33Apr 9, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Artifact for "Marconi: Prefix Caching for the Era of Hybrid LLMs" [MLSys '25 Outstanding Paper Award, Honorable Mention]☆66Mar 5, 2025Updated last year
- ⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.☆158May 10, 2025Updated last year
- Examples of CUDA implementations by Cutlass CuTe☆281Jul 1, 2025Updated last year
- Code for "Fast Sparse ConvNets" CVPR2020 submissions☆12Nov 20, 2019Updated 6 years ago
- DeepSeek-V3/R1 inference performance simulator☆194Mar 27, 2025Updated last year
- ☆100Apr 2, 2025Updated last year
- Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of pap…☆284Mar 6, 2025Updated last year
- Dynamic Memory Management for Serving LLMs without PagedAttention☆512Jul 17, 2026Updated 3 weeks ago
- ☆37Aug 7, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,629Updated this week
- Sequence-level 1F1B schedule for LLMs.☆37Aug 26, 2025Updated 11 months ago
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆546Updated this week
- ☆49Apr 15, 2024Updated 2 years ago
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆75Updated this week
- Ventus GPGPU ISA Simulator Based on Spike☆53Updated this week
- ☆186Updated this week