Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotuning, and advanced scheduling. Run benchmarks across different GPU types and configurations, track results and share experiments with the community.
☆79Aug 23, 2026Updated this week
Alternatives and similar repositories for emmy
Users that are interested in emmy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- 。☆13Jan 15, 2022Updated 4 years ago
- Arithmetic multiplier benchmarks☆12Nov 13, 2017Updated 8 years ago
- F4 algorithm C++ library (groebner basis computations over finite fields)☆14Apr 6, 2018Updated 8 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- High-Performance Structured Linear Operators☆13May 17, 2018Updated 8 years ago
- ☆34Jul 30, 2026Updated 3 weeks ago
- For GL-iNET SF1200/SFT1200☆18Mar 15, 2024Updated 2 years ago
- Implementation of local search-based algorithms for solving SAT and Max-SAT in Python☆13Dec 6, 2020Updated 5 years ago
- A fork of the Kissat SAT solver with additional features. Supports incremental solving.☆17Aug 13, 2022Updated 4 years ago
- Make FP4 on 5090 Great Again☆22Jul 20, 2026Updated last month
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- Tenstorrent Topology (TT-Topology) is a command line utility used to flash multiple NB cards on a system to use specific eth routing conf…☆16Jul 23, 2026Updated last month
- ☆10Aug 14, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A design automation framework to engineer decision diagrams yourself☆27Aug 11, 2026Updated last week
- LLFree: Lock- and Log-free Allocator☆12Jun 8, 2026Updated 2 months ago
- A modular Lustre to C / Horn clauses compiler☆22Nov 17, 2018Updated 7 years ago
- Benchmarking Intelligence Efficiency of LM Inference☆85Updated this week
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆109Jul 21, 2026Updated last month
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆40Jul 22, 2026Updated last month
- High-performance GPU kernels for Ads and Recsys model training, independently implemented and optimized for real-world workloads and mode…☆37Aug 7, 2026Updated 2 weeks ago
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆18Jul 26, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A blazing fast lightweight static compile time 2D, 3D, and N-dimensional C++20 vector library.☆25Dec 12, 2023Updated 2 years ago
- Repository for the CUDA H100 Course☆73Apr 12, 2026Updated 4 months ago
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆30Dec 17, 2024Updated last year
- Collection of official scripts created by the Dione Team.☆15Feb 21, 2026Updated 6 months ago
- Volume Manipulation Library☆17Jul 13, 2023Updated 3 years ago
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 6 months ago
- Mistral Vibe rewritten in Rust by Devstral 2☆21Dec 23, 2025Updated 8 months ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆59Feb 25, 2026Updated 5 months ago
- Web frontend for Myria☆12Sep 30, 2020Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 6 months ago
- The shared memory version of the Alternating Directions Implicit Solver for Isogeometric Analysis☆10Jan 26, 2019Updated 7 years ago
- ☆11Jun 11, 2020Updated 6 years ago
- Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!☆301Jul 26, 2026Updated 3 weeks ago
- How to create and record demos in terminal sessions☆11May 3, 2024Updated 2 years ago
- In-process, multi master, distributed database☆21Jun 18, 2026Updated 2 months ago
- GPEmu, a GPU emulator for faster and cheaper prototyping and evaluation of deep learning system research☆45Dec 2, 2024Updated last year