Artifact from "Hardware Compute Partitioning on NVIDIA GPUs". THIS IS A FORK OF BAKITAS REPO. I AM NOT ONE OF THE AUTHORS OF THE PAPER.
☆67Nov 24, 2025Updated 8 months ago
Alternatives and similar repositories for libsmctrl
Users that are interested in libsmctrl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆33Jul 13, 2026Updated last month
- ☆27Aug 19, 2022Updated 3 years ago
- Tutorials for NVIDIA CUPTI samples☆72Jul 22, 2026Updated 3 weeks ago
- ☆21Jul 7, 2026Updated last month
- An interference-aware scheduler for fine-grained GPU sharing☆164Nov 26, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- a simple API to use CUPTI☆10Aug 19, 2025Updated 11 months ago
- Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.☆74Mar 17, 2026Updated 4 months ago
- A GPU-accelerated DNN inference serving system that supports instant kernel preemption and biased concurrent execution in GPU scheduling.☆43May 29, 2022Updated 4 years ago
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆75Updated this week
- ☆12Nov 5, 2024Updated last year
- ☆267Dec 25, 2025Updated 7 months ago
- ☆34Sep 9, 2020Updated 5 years ago
- A highly-flexible GPU simulator for AMD GPUs.☆263Aug 5, 2026Updated last week
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond☆1,131Updated this week
- ☆12Aug 17, 2022Updated 3 years ago
- eBPF for GPU UVM offloading and scheduling in Linux kernel☆63Apr 15, 2026Updated 3 months ago
- ☆25May 18, 2025Updated last year
- ☆79May 4, 2021Updated 5 years ago
- Automatic Parallelism Using LLVM☆10Aug 2, 2014Updated 12 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆47May 13, 2025Updated last year
- collection of benchmarks to measure basic GPU capabilities☆533Oct 24, 2025Updated 9 months ago
- [ACM EuroSys 2023] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access☆56Aug 6, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Efficient and easy multi-instance LLM serving☆563Mar 12, 2026Updated 5 months ago
- Experiments evaluating preemption on the NVIDIA Pascal architecture☆16Nov 10, 2016Updated 9 years ago
- Allow torch tensor memory to be released and resumed later☆267Updated this week
- Dynamic Memory Management for Serving LLMs without PagedAttention☆512Jul 17, 2026Updated 3 weeks ago
- Hooked CUDA-related dynamic libraries by using automated code generation tools.☆173Dec 12, 2023Updated 2 years ago
- An efficient storage system for concurrent graph processing☆10Feb 1, 2021Updated 5 years ago
- [ASPLOS' 26] TetriServe: Efficiently Serving Mixed DiT Workloads☆17Mar 12, 2026Updated 5 months ago
- ☆85Apr 18, 2025Updated last year
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoS☆33Feb 10, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Unofficial description of the CUDA assembly (SASS) instruction sets.☆238Jul 18, 2025Updated last year
- Artifacts for our NSDI'23 paper TGS☆99Jun 10, 2024Updated 2 years ago
- ☆18Apr 21, 2024Updated 2 years ago
- CUDA checkpoint and restore utility☆481Jul 6, 2026Updated last month
- Evaluation utilities based on SymPy.☆25Dec 12, 2024Updated last year
- A NCCL extension library, designed to efficiently offload GPU memory allocated by the NCCL communication library.☆115Dec 17, 2025Updated 7 months ago
- Assembler and Decompiler for NVIDIA (Maxwell Pascal Volta Turing Ampere) GPUs.☆97Feb 23, 2023Updated 3 years ago