Hardware acceleration for transformer attention mechanisms on NVIDIA Deep Learning Accelerator (NVDLA), enabling efficient inference of transformer models with 135× better power efficiency than CPU implementations.
☆17Mar 30, 2025Updated last year
Alternatives and similar repositories for nvdla-attn-mechanism
Users that are interested in nvdla-attn-mechanism are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Anatomy of a powerhouse: SystemVerilog TPU based on Google TPU v1☆23Nov 9, 2025Updated 8 months ago
- Linux on RISC-V on FPGA (LOROF): RV64GC Sv39 Quad-Core Superscalar Out-of-Order Virtual Memory CPU☆18Updated this week
- Visualization tool for designing mesh Network-on-Chips (NoC) and assisting with architecture research☆17Jan 21, 2024Updated 2 years ago
- RTL implementation of a ray-tracing GPU☆16Dec 18, 2012Updated 13 years ago
- RISC-V vector and tensor compute extensions for Vortex GPGPU acceleration for ML workloads. Optimized for transformer models, CNNs, and g…☆25Apr 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- FSA: Fusing FlashAttention within a Single Systolic Array☆186Apr 15, 2026Updated 3 months ago
- This is a project created and completed by team BOOM(Beihang OO masters).This is a superscalar processor with a 13-stage out-of-order dua…☆18Sep 29, 2024Updated last year
- Template for project1 TPU☆23May 1, 2021Updated 5 years ago
- RISC-V SIMD Superscalar Dual-Issue Processor☆31Apr 24, 2025Updated last year
- ☆79Apr 22, 2025Updated last year
- Open-source AI Accelerator Stack integrating compute, memory, and software — from RTL to PyTorch.☆26Jul 2, 2026Updated 2 weeks ago
- Learn NVDLA by SOMNIA☆42Dec 13, 2019Updated 6 years ago
- Gem5 with chinese comment and introduction (master) and some other std gem5 version.☆43Jan 2, 2022Updated 4 years ago
- RTL code for the DPU chip designed for irregular graphs☆14May 30, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Course Project for High Level Chip Design (高层次芯片设计)☆18Jan 2, 2025Updated last year
- 给NEMU移植Linux Kernel!☆23Jun 1, 2025Updated last year
- Simulator of a memory controller to connect DRAMSim and FlashDIMMSim into one unified memory☆18Apr 4, 2024Updated 2 years ago
- ☆37Sep 30, 2025Updated 9 months ago
- A co-simulation framework for chiplet-based systems executing DNN models.☆17Feb 16, 2026Updated 5 months ago
- [DATE 2025] Official implementation and dataset of AIrchitect v2: Learning the Hardware Accelerator Design Space through Unified Represen…☆20Jan 17, 2025Updated last year
- Infrastructure to drive Spike (RISC-V ISA Simulator) in cosim mode. Hammer provides a C++ and Python interface to interact with Spike.☆39Jan 19, 2026Updated 6 months ago
- GPGPU-Sim 中文注释版代码,包含 GPGPU-Sim 模拟器的最新版代码,经过中文注释,以帮助中文用户更好地理解和使用该模拟器。☆30Dec 18, 2024Updated last year
- OpenTestability is an open-source tool for structural analysis of digital circuits, enabling computation of SCOAP metrics, Controllibilit…☆20Jul 13, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆17Sep 15, 2023Updated 2 years ago
- Here are some implementations of basic hardware units in RTL language (verilog for now), which can be used for area/power evaluation and …☆15Aug 25, 2023Updated 2 years ago
- Custom ASIC Design for SHA-256☆14Nov 22, 2025Updated 7 months ago
- This is the repository containing the implementation of sparse dense matrix multiplication for the matrix dimension of 560 x 560.☆10Jul 7, 2021Updated 5 years ago
- Petri Net Simulator program☆10Nov 27, 2017Updated 8 years ago
- simple RISC-V 64bit emulator, which can boot linux kernel.☆12Oct 16, 2023Updated 2 years ago
- Artifact material for [HPCA 2025] #2108 "UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing Architectures"☆60Sep 1, 2025Updated 10 months ago
- AXI Verification using UVM Testbench☆16Oct 20, 2024Updated last year
- Learn and build GPU RTL from scratch☆22Aug 1, 2025Updated 11 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆49Nov 18, 2019Updated 6 years ago
- A custom AI chip to be taped out soon!☆49Dec 20, 2025Updated 7 months ago
- C++ code for HLS FPGA implementation of transformer☆24Sep 11, 2024Updated last year
- Final year research project to design a programmable virtual switch based on the specifications of a TSN to be implemented on a TSN netwo…☆13Nov 17, 2020Updated 5 years ago
- Hardware implementation of a Fixed Point Recursive Forward and Inverse FFT algorithm☆17Mar 3, 2018Updated 8 years ago
- TeLLMe: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs [FPGA2026]☆32Mar 11, 2026Updated 4 months ago
- ☆20Nov 18, 2022Updated 3 years ago