Benchmark and deploy optimized LLM models on GPU servers with vLLM or SGLang. Chose from a list of optimized recipes for popular models or create your own with custom configurations. Run benchmarks across different GPU types and configurations, track results, and share experiments with the community.
☆67Jul 28, 2026Updated this week
Alternatives and similar repositories for emmy
Users that are interested in emmy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Arithmetic multiplier benchmarks☆12Nov 13, 2017Updated 8 years ago
- Decentralizing distribution of open-source AI models.☆16Updated this week
- F4 algorithm C++ library (groebner basis computations over finite fields)☆14Apr 6, 2018Updated 8 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- High-Performance Structured Linear Operators☆13May 17, 2018Updated 8 years ago
- For GL-iNET SF1200/SFT1200☆18Mar 15, 2024Updated 2 years ago
- Make FP4 on 5090 Great Again☆17Jul 20, 2026Updated last week
- Parsing library for BLIF netlists☆19Nov 1, 2024Updated last year
- Pure C wrapper library to use llama.cpp with Linux and Windows as simple as possible.☆15Updated this week
- Source code and dataset for the paper 'Saamayik: A Benchmark and Dataset for English-Sanskrit Translation'☆16Oct 11, 2025Updated 9 months ago
- Programming language w/ subproject that implements the Go scheduler in C++☆12Jan 21, 2018Updated 8 years ago
- A collection of some awesome public MAX platform, Mojo programming language and Multi-Level IR Compiler Framework(MLIR) projects.☆44Dec 28, 2024Updated last year
- A design automation framework to engineer decision diagrams yourself☆27Jul 19, 2026Updated last week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Lightweight Python Wrapper for OpenVINO, enabling LLM inference on NPUs☆29Dec 17, 2024Updated last year
- axum_embed is a library that provides a service for serving embedded files using the axum web framework.☆20Jan 6, 2025Updated last year
- A code sample demonstrating how to share and rebuild a PyTorch GPU tensor via its pointer/reference between different processes.☆15Aug 27, 2024Updated last year
- Benchmarking Intelligence Efficiency of LM Inference☆72Jul 17, 2026Updated last week
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold (NeurIPS 2025 Spotlight)☆17Jun 29, 2026Updated 3 weeks ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆33Jul 22, 2026Updated last week
- High-performance GPU kernels for Ads and Recsys model training, independently implemented and optimized for real-world workloads and mode…☆34Jul 10, 2026Updated 2 weeks ago
- Repository for the CUDA H100 Course☆67Apr 12, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Collection of official scripts created by the Dione Team.☆15Feb 21, 2026Updated 5 months ago
- Volume Manipulation Library☆17Jul 13, 2023Updated 3 years ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆115Jun 28, 2025Updated last year
- ☆13Jul 10, 2026Updated 2 weeks ago
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 5 months ago
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆78Jul 12, 2026Updated 2 weeks ago
- Mistral Vibe rewritten in Rust by Devstral 2☆20Dec 23, 2025Updated 7 months ago
- A simple SAT solver based on the CDCL algorithm☆20Sep 10, 2019Updated 6 years ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆56Feb 25, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- MonetDB driver for Go☆11Apr 24, 2018Updated 8 years ago
- This repository contains scripts and guides presented during multiple LLM development sessions by me.☆22Apr 5, 2026Updated 3 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆79Feb 18, 2026Updated 5 months ago
- Local LLM Server Manager + LlaMA.cpp + Chat☆17Updated this week
- Go binding for libshout 2.x☆13Oct 3, 2013Updated 12 years ago
- The shared memory version of the Alternating Directions Implicit Solver for Isogeometric Analysis☆10Jan 26, 2019Updated 7 years ago
- Lightweight Linux namespace utilities☆17Jul 27, 2019Updated 7 years ago