Autonomous CUDA kernel optimization agent with structured task specs and per-config scoring
☆17Jun 17, 2026Updated 3 months ago
Alternatives and similar repositories for ferret
Users that are interested in ferret are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An Attention Superoptimizer☆22Jan 20, 2025Updated last year
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆23Nov 28, 2025Updated 10 months ago
- Compact and Agent-Native MoE Training System☆355Sep 23, 2026Updated last week
- A mini version of k8s that implements the abstraction of pod, service, auto-scaling, replicaSet and provides DNS, GPU and serverless serv…☆18Jun 16, 2023Updated 3 years ago
- Group project of SE3356 Cloud Operating System Design and Practice, Spring 2022.☆24Jun 21, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 4 months ago
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- mKernel: fast multi-node, multi-GPU fused kernels☆286Updated this week
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- [NSDI'26] PolyRL is a reinforcement learning framework for LLM that harvest spot instances on the cloud to reduce cost.☆19Mar 30, 2026Updated 6 months ago
- Artifact of Chimera☆18May 6, 2025Updated last year
- ☆36Jul 19, 2025Updated last year
- [EuroSys'25] Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization☆24Apr 13, 2026Updated 5 months ago
- Website for CSE 234, Winter 2025☆16Mar 24, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A Homework for Computer Architecture at SJTU☆14Jan 4, 2020Updated 6 years ago
- OSDI 2023 Welder, deeplearning compiler☆36Nov 24, 2023Updated 2 years ago
- Virtual Decoupled Cores: Composable Programming Framework and Runtime for Async GPUs☆23Updated this week
- [Archived] For the latest updates and community contribution, please visit: https://github.com/Ascend/TransferQueue or https://gitcode.co…☆15Aug 14, 2026Updated last month
- ☆25Jun 12, 2023Updated 3 years ago
- paper and its code for AI System☆379May 14, 2026Updated 4 months ago
- ☆19May 9, 2025Updated last year
- ☆46Oct 15, 2025Updated 11 months ago
- ☆12May 18, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A home for the final text of all TVM RFCs.☆111Sep 24, 2024Updated 2 years ago
- A simple containerized application manage system like Kubernetes, but written in Rust☆19Jun 25, 2022Updated 4 years ago
- ☆26Jun 10, 2026Updated 3 months ago
- ☆54Aug 6, 2024Updated 2 years ago
- ☆21Mar 17, 2026Updated 6 months ago
- 上海交通大学软件学院课程计算机系统基础(ICS)笔记☆15Feb 7, 2022Updated 4 years ago
- ☆51Aug 17, 2026Updated last month
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- Collection of scripts to build PyTorch and the domain libraries from source.☆14Sep 18, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- a size profiler for cuda binary☆73Updated this week
- ☆27Oct 1, 2025Updated last year
- ⚡ Bring some magic to i.sjtu.edu.cn☆22Jan 3, 2020Updated 6 years ago
- [ICML 2025] SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity☆79Mar 10, 2026Updated 6 months ago
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designers☆69Aug 22, 2026Updated last month
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆105Apr 7, 2026Updated 5 months ago
- ☆33Jul 17, 2024Updated 2 years ago