An agent for CUDA compute-communication kernel co-design
☆35May 7, 2026Updated 2 months ago
Alternatives and similar repositories for cuco
Users that are interested in cuco are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Jun 17, 2026Updated last month
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆35May 22, 2026Updated last month
- Lightweight agent multiplexer, all in one Web dashboard☆50Updated this week
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- [NSDI'26] PolyRL is a reinforcement learning framework for LLM that harvest spot instances on the cloud to reduce cost.☆19Mar 30, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆21Nov 28, 2025Updated 7 months ago
- [ASPLOS' 26] TetriServe: Efficiently Serving Mixed DiT Workloads☆17Mar 12, 2026Updated 4 months ago
- Automated High-Performance GPU Kernel Generation☆120Jun 1, 2026Updated last month
- ☆45Oct 15, 2025Updated 9 months ago
- ☆19May 9, 2025Updated last year
- A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search☆21Jul 22, 2025Updated 11 months ago
- Autonomous CUDA kernel optimization agent with structured task specs and per-config scoring☆17Jun 17, 2026Updated last month
- a simple API to use CUPTI☆10Aug 19, 2025Updated 11 months ago
- ☆21Jun 6, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [Accepted to SOSP 2026] Fast Deterministic LLM Inference☆22Updated this week
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆106Apr 7, 2026Updated 3 months ago
- ☆37Aug 7, 2025Updated 11 months ago
- ☆56Aug 27, 2024Updated last year
- ML kernels and benchmarking infrastructure written in TIRx☆66Updated this week
- diffusers with search engine☆12Jan 13, 2026Updated 6 months ago
- [Archived] For the latest updates and community contribution, please visit: https://github.com/Ascend/TransferQueue or https://gitcode.co…☆16Jan 16, 2026Updated 6 months ago
- A Triton-only attention backend for vLLM☆27Updated this week
- ☆28Nov 29, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- DeDe (OSDI '25): an optimization framework for large-scale resource allocation☆15May 18, 2026Updated 2 months ago
- Can AI Agents Build Bespoke Systems?☆84Updated this week
- A framework and CLI toolkit for orchestrating teams of loosely-coupled AI agents.☆18Updated this week
- Fast, memory-efficient attention column reduction (e.g., sum, mean, max)☆49Feb 10, 2026Updated 5 months ago
- ☆41Dec 9, 2025Updated 7 months ago
- Artifacts of EVT ASPLOS'24☆29Mar 6, 2024Updated 2 years ago
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- Implementation from scratch in C of the Multi-head latent attention used in the Deepseek-v3 technical paper.☆18Jan 15, 2025Updated last year
- ☆140Feb 17, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Chunky Loop Interaction☆25Aug 13, 2019Updated 6 years ago
- Global Address SPace toolbox -- Julia wrapper☆10Nov 17, 2017Updated 8 years ago
- Modular simulation library for AI datacenter-grid interaction☆15May 11, 2026Updated 2 months ago
- RPCNIC: A High-Performance and Reconfigurable PCIe-attached RPC Accelerator [HPCA2025]☆15Dec 9, 2024Updated last year
- ☆40Dec 14, 2025Updated 7 months ago
- Welcome to PeriFlow CLI ☁︎☆12Aug 3, 2023Updated 2 years ago
- ☆33Mar 12, 2026Updated 4 months ago