A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
☆80Aug 27, 2026Updated this week
Alternatives and similar repositories for Magpie
Users that are interested in Magpie are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Automated bottleneck detection and solution orchestration☆23Feb 24, 2026Updated 6 months ago
- Generating Efficient AI-Centric Kernels☆163Updated this week
- Repo containing artifacts for Neurips 2025 tutorial- How to Build Agents to Generate Kernels for Faster LLMs (and Other Models!)☆15May 11, 2026Updated 3 months ago
- Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profile…☆76Updated this week
- Automating analysis from trace files☆90Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆30Aug 24, 2026Updated last week
- Simply log all kernel durations☆18Updated this week
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- generate music visualizations in the browser with threejs☆37Oct 5, 2025Updated 10 months ago
- The C++ Standard Library for your entire system.☆28Updated this week
- Fast and Furious AMD Kernels☆461Updated this week
- ☆10May 15, 2024Updated 2 years ago
- ☆71Updated this week
- A ROCm library for GPU-Initiated IO. This provides support for initiating IO from a ROCm-capable GPU against a range of targets including…☆57Aug 8, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated last month
- Header-only library of GPU-accelerated, concurrent data structures.☆12Jul 10, 2026Updated last month
- IntelliKit is a collection of intelligent tools designed to make GPU kernel development, profiling, and validation accessible to LLMs and…☆32Updated this week
- AMD 0.9B efficient text to video diffusion model☆48May 16, 2026Updated 3 months ago
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated last month
- AiTer Optimized Model☆168Updated this week
- HRX: Hip Runtime Extended☆27Updated this week
- ☆33Oct 2, 2025Updated 10 months ago
- Code repository dedicated to experimenting and research with tiny reasoning language model☆52Nov 24, 2025Updated 9 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- FlashSampling: Fast and Memory-Efficient Exact Sampling (https://huggingface.co/papers/2603.15854)☆87Aug 5, 2026Updated 3 weeks ago
- ☆25Sep 12, 2025Updated 11 months ago
- Utilities for ROCm Tech Support Log Collections☆14May 29, 2026Updated 3 months ago
- Minimal Implimentation of VCRec (2024) for collapse provention.☆18Jan 28, 2025Updated last year
- Production-ready ternary quantized (1.58-bit) Rust code generation model with mHC-lite, MaxRL training, and comprehensive benchmarking☆22Aug 2, 2026Updated 3 weeks ago
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆121Updated this week
- ☆31May 10, 2026Updated 3 months ago
- GPU kernel benchmarking☆47Jun 10, 2026Updated 2 months ago
- Started as a Team Project for CS690D at UMass Amherst, now turning into pytorch implementation of hyperbolic neural networks using Poinca…☆12Dec 8, 2022Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆15Jan 27, 2025Updated last year
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆25May 5, 2026Updated 3 months ago
- HIP backend patch for Numba, the NumPy aware dynamic Python compiler using LLVM.☆22Jul 10, 2026Updated last month
- ☆17Aug 11, 2025Updated last year
- Multi-Turn RL Training System with AgentTrainer for Language Model Game Reinforcement Learning☆67Dec 18, 2025Updated 8 months ago
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆198Updated this week
- vibe coding some meta humans☆24May 25, 2026Updated 3 months ago