Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotuning, and advanced scheduling. Run benchmarks across different GPU types and configurations, track results and share experiments with the community.
☆80Sep 12, 2026Updated this week
Alternatives and similar repositories for emmy
Users that are interested in emmy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 7 months ago
- ☆36Updated this week
- Implementation of local search-based algorithms for solving SAT and Max-SAT in Python☆13Dec 6, 2020Updated 5 years ago
- A fork of the Kissat SAT solver with additional features. Supports incremental solving.☆17Aug 13, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Minimize server usage by leveraging a decentralized peer-to-peer network for ultra-low-latency live streaming among users.☆13Feb 19, 2024Updated 2 years ago
- Make FP4 on 5090 Great Again☆24Jul 20, 2026Updated last month
- Parsing library for BLIF netlists☆19Nov 1, 2024Updated last year
- Pure C wrapper library to use llama.cpp with Linux and Windows as simple as possible.☆15Sep 6, 2026Updated last week
- Source code and dataset for the paper 'Saamayik: A Benchmark and Dataset for English-Sanskrit Translation'☆16Oct 11, 2025Updated 11 months ago
- axum_embed is a library that provides a service for serving embedded files using the axum web framework.☆20Jan 6, 2025Updated last year
- Benchmarking Intelligence Efficiency of LM Inference☆87Updated this week
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆43Jul 22, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- High-performance GPU kernels for Ads and Recsys model training, independently implemented and optimized for real-world workloads and mode…☆42Aug 25, 2026Updated 2 weeks ago
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆30Dec 17, 2024Updated last year
- Collection of official scripts created by the Dione Team.☆15Feb 21, 2026Updated 6 months ago
- Volume Manipulation Library☆16Jul 13, 2023Updated 3 years ago
- Repository for the CUDA H100 Course☆78Apr 12, 2026Updated 5 months ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆118Aug 4, 2026Updated last month
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 6 months ago
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆84Jul 12, 2026Updated 2 months ago
- Quantize transformers to any learned arbitrary 4-bit numeric format☆59Jul 2, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆21Jul 26, 2026Updated last month
- Web frontend for Myria☆12Sep 30, 2020Updated 5 years ago
- This repository contains scripts and guides presented during multiple LLM development sessions by me.☆22Apr 5, 2026Updated 5 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆82Feb 18, 2026Updated 6 months ago
- The shared memory version of the Alternating Directions Implicit Solver for Isogeometric Analysis☆10Jan 26, 2019Updated 7 years ago
- ☆11Jun 11, 2020Updated 6 years ago
- Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!☆304Jul 26, 2026Updated last month
- GPEmu, a GPU emulator for faster and cheaper prototyping and evaluation of deep learning system research☆45Dec 2, 2024Updated last year
- Local LLM Server Manager + LlaMA.cpp + Chat☆135Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated last month
- A toy nanopass compiler for x86 written in lean☆15Oct 25, 2025Updated 10 months ago
- ☆30Jan 26, 2023Updated 3 years ago
- A simple single-threaded concurrency runtime for Rust based on io_uring.☆27Jan 6, 2024Updated 2 years ago
- Type inference implementation in OCaml using Algorithm W☆10Aug 26, 2021Updated 5 years ago
- The AI-Native SDLC Playbook — Anthropic Claude Academy course as EPUB ebook☆68Sep 3, 2026Updated last week
- PIDX☆14Jan 20, 2020Updated 6 years ago