Lynn 原生 LLM 推理引擎 · W4A8/NVFP4 量化 · 自写 CUDA/Triton kernel · MoE · 投机解码 | Lynn-native LLM inference engine for NVIDIA Blackwell
☆27Jun 8, 2026Updated 3 months ago
Alternatives and similar repositories for lynn-engine
Users that are interested in lynn-engine are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Jun 16, 2026Updated 2 months ago
- 一个轻量化的大模型推理框架☆23May 26, 2025Updated last year
- [Accepted to SOSP 2026] Fast Deterministic LLM Inference☆28Updated this week
- Running LLaMA 3 with Rust.☆10May 21, 2024Updated 2 years ago
- ☆11Jan 11, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- library which simplifies host-GPU data transfer using userspace pagefault handling☆15Jun 8, 2012Updated 14 years ago
- Common template for pytorch project. Easy to extent and modify for new project.☆13Dec 13, 2022Updated 3 years ago
- ☆14Mar 1, 2020Updated 6 years ago
- Sampling techniques for Candle.☆21Apr 3, 2024Updated 2 years ago
- This repo models rowhammer in gem5.☆11Aug 7, 2026Updated last month
- Java-like Language with Static Information Flow Types☆16May 5, 2025Updated last year
- A source-to-source compiler for optimizing CUDA dynamic parallelism by aggregating launches☆15Jun 21, 2019Updated 7 years ago
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago
- ☆18Mar 19, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A Feishu/Lark AI agent bot☆15Feb 27, 2026Updated 6 months ago
- ☆20Oct 5, 2025Updated 11 months ago
- A straightforward (complete) sample of how to implement AES-GCM by using Linux crypto API at kernel side☆13Oct 6, 2022Updated 3 years ago
- ☆19May 30, 2024Updated 2 years ago
- ☆14Jan 12, 2022Updated 4 years ago
- ☆14Mar 7, 2025Updated last year
- ☆13Oct 6, 2024Updated last year
- A high-performance C/C++ inference server for Qwen3-ASR , optimized for CPU/GPU real-time streaming speech recognition.☆17Jun 27, 2026Updated 2 months ago
- GPU Power Modelling Tool☆14Nov 15, 2019Updated 6 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Eagle and EagleSim: Deep-RL for PTZ Cameras☆10Aug 23, 2024Updated 2 years ago
- ☆17May 22, 2023Updated 3 years ago
- cutile kernel examples☆51Apr 3, 2026Updated 5 months ago
- Keep a journal of things I want to share☆12Aug 23, 2022Updated 4 years ago
- Rust standalone inference of Namo-500M series models. Extremly tiny, runing VLM on CPU.☆24Mar 12, 2025Updated last year
- Simulator of a memory controller to connect DRAMSim and FlashDIMMSim into one unified memory☆18Apr 4, 2024Updated 2 years ago
- Candle Pipelines provides a simple, intuitive interface for Rust developers who want to work with Large Language Models locally, powered …☆23Jan 5, 2026Updated 8 months ago
- safe and simple append only database written in C☆23Sep 6, 2016Updated 10 years ago
- ☆22Aug 28, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Translate Virtual Address To Physical Address in Linux Kernel☆17Dec 27, 2019Updated 6 years ago
- ☆13Feb 9, 2026Updated 7 months ago
- This repository is for sharing code and information related to researching the "rowhammer" problem with respect to GPUs.☆15Apr 20, 2017Updated 9 years ago
- ☆17Sep 20, 2021Updated 4 years ago
- TensorRT for RefineNet Segmentation☆12Apr 27, 2021Updated 5 years ago
- ☆15May 8, 2025Updated last year
- Flash Attention in ~100 lines of CUDA (forward pass only)☆12Jun 10, 2024Updated 2 years ago