Artifact from "Hardware Compute Partitioning on NVIDIA GPUs". THIS IS A FORK OF BAKITAS REPO. I AM NOT ONE OF THE AUTHORS OF THE PAPER.
☆67Nov 24, 2025Updated 9 months ago
Alternatives and similar repositories for libsmctrl
Users that are interested in libsmctrl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34Jul 13, 2026Updated last month
- ☆27Aug 19, 2022Updated 4 years ago
- Tutorials for NVIDIA CUPTI samples☆72Jul 22, 2026Updated last month
- ☆22Updated this week
- An interference-aware scheduler for fine-grained GPU sharing☆163Nov 26, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- a simple API to use CUPTI☆10Aug 19, 2025Updated last year
- Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.☆77Mar 17, 2026Updated 5 months ago
- A GPU-accelerated DNN inference serving system that supports instant kernel preemption and biased concurrent execution in GPU scheduling.☆43May 29, 2022Updated 4 years ago
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆76Aug 8, 2026Updated 3 weeks ago
- ☆12Nov 5, 2024Updated last year
- ☆269Dec 25, 2025Updated 8 months ago
- ☆34Sep 9, 2020Updated 5 years ago
- A highly-flexible GPU simulator for AMD GPUs.☆264Updated this week
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,276Aug 23, 2026Updated last week
- eBPF for GPU UVM offloading and scheduling in Linux kernel☆67Updated this week
- ☆25May 18, 2025Updated last year
- ☆79May 4, 2021Updated 5 years ago
- Automatic Parallelism Using LLVM☆10Aug 2, 2014Updated 12 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆49May 13, 2025Updated last year
- collection of benchmarks to measure basic GPU capabilities☆538Oct 24, 2025Updated 10 months ago
- [ACM EuroSys 2023] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access☆56Aug 6, 2025Updated last year
- Efficient and easy multi-instance LLM serving☆563Mar 12, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Experiments evaluating preemption on the NVIDIA Pascal architecture☆16Nov 10, 2016Updated 9 years ago
- Allow torch tensor memory to be released and resumed later☆273Aug 21, 2026Updated last week
- Dynamic Memory Management for Serving LLMs without PagedAttention☆519Aug 24, 2026Updated last week
- GeminiFS: A Companion File System for GPUs☆86Aug 11, 2026Updated 3 weeks ago
- Hooked CUDA-related dynamic libraries by using automated code generation tools.☆173Dec 12, 2023Updated 2 years ago
- An efficient storage system for concurrent graph processing☆10Feb 1, 2021Updated 5 years ago
- [ASPLOS' 26] TetriServe: Efficiently Serving Mixed DiT Workloads☆18Mar 12, 2026Updated 5 months ago
- ☆87Apr 18, 2025Updated last year
- Unofficial description of the CUDA assembly (SASS) instruction sets.☆244Jul 18, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Artifacts for our NSDI'23 paper TGS☆98Jun 10, 2024Updated 2 years ago
- ☆18Apr 21, 2024Updated 2 years ago
- CUDA checkpoint and restore utility☆485Jul 6, 2026Updated last month
- Evaluation utilities based on SymPy.☆25Dec 12, 2024Updated last year
- A NCCL extension library, designed to efficiently offload GPU memory allocated by the NCCL communication library.☆117Dec 17, 2025Updated 8 months ago
- Assembler and Decompiler for NVIDIA (Maxwell Pascal Volta Turing Ampere) GPUs.☆98Feb 23, 2023Updated 3 years ago
- A group of students who are interested in Compilers, and they want to improve themselves together.☆24Aug 23, 2022Updated 4 years ago