A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1
☆1,373Apr 3, 2026Updated 4 months ago
Alternatives and similar repositories for tiny-tpu
Users that are interested in tiny-tpu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A machine learning accelerator core designed for energy-efficient AI at the edge.☆2,539Updated this week
- ☆174Jan 4, 2026Updated 7 months ago
- a mini 2x2 systolic array and PE demo☆75Dec 21, 2025Updated 8 months ago
- A minimal GPU design in Verilog to learn how GPUs work from the ground up☆12,901Aug 18, 2024Updated 2 years ago
- A custom AI chip to be taped out soon!☆50Dec 20, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- FSA: Fusing FlashAttention within a Single Systolic Array☆196Apr 15, 2026Updated 4 months ago
- A minimal Tensor Processing Unit (TPU) inspired by Google's TPUv1.☆205Aug 10, 2024Updated 2 years ago
- Berkeley's Spatial Array Generator☆1,443Updated this week
- minimal compiler☆24Feb 19, 2026Updated 6 months ago
- ☆2,231Updated this week
- opensource NPU for LLM inference (this run gpt2)☆239Feb 16, 2026Updated 6 months ago
- Anatomy of a powerhouse: SystemVerilog TPU based on Google TPU v1☆25Nov 9, 2025Updated 9 months ago
- A open source reimplementation of Google's Tensor Processing Unit (TPU).☆778Dec 6, 2017Updated 8 years ago
- An AI accelerator implementation with Xilinx FPGA☆123Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Hardware acceleration for transformer attention mechanisms on NVIDIA Deep Learning Accelerator (NVDLA), enabling efficient inference of…☆20Mar 30, 2025Updated last year
- ☆87Apr 22, 2025Updated last year
- RV64GC Linux Capable RISC-V Core☆69Aug 7, 2026Updated 3 weeks ago
- Curriculum for a university course to teach chip design using open source EDA tools☆163Oct 21, 2023Updated 2 years ago
- A PULP SoC for education, easy to understand and extend with a full flow for a physical design.☆270Updated this week
- ONNXim is a fast cycle-level simulator that can model multi-core NPUs for DNN inference☆210Jan 8, 2026Updated 7 months ago
- SystemVerilog Implementations of CUDA/TensorCore/TPU GEMM Operations☆21Apr 12, 2026Updated 4 months ago
- a mini TPU with floating point arithmetic☆53Dec 22, 2025Updated 8 months ago
- Implementation of a Tensor Processing Unit for embedded systems and the IoT.☆575Jan 5, 2019Updated 7 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Linux on RISC-V on FPGA (LOROF): RV64GC Sv39 Quad-Core Superscalar Out-of-Order Virtual Memory CPU☆19Updated this week
- CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-base…☆1,020Updated this week
- OpenSource GPU, in Verilog, loosely based on RISC-V ISA☆1,366Nov 22, 2024Updated last year
- Small-scale Tensor Processing Unit built on an FPGA☆230Aug 4, 2019Updated 7 years ago
- Fork of github.com/UCSBarchlab/OpenTPU for the TGPTPU project☆15Jun 1, 2025Updated last year
- CORE-V Wally is a configurable RISC-V Processor associated with RISC-V System-on-Chip Design textbook. Contains a 5-stage pipeline, suppo…☆611Updated this week
- Spatz is a compact RISC-V-based vector processor meant for high-performance, small computing clusters.☆169Updated this week
- Allo Accelerator Design and Programming Framework (PLDI'24)☆409Updated this week
- Open-source high-performance RISC-V processor☆7,234Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Single Instruction Multiple Threads GPU Core with textbook Streaming Multi-Processor features☆71Jan 30, 2026Updated 7 months ago
- SAURIA (Systolic-Array tensor Unit for aRtificial Intelligence Acceleration) is an open-source Convolutional Neural Network accelerator b…☆116Nov 26, 2025Updated 9 months ago
- Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.☆490Aug 17, 2026Updated 2 weeks ago
- An Agile RISC-V SoC Design Framework with in-order cores, out-of-order cores, accelerators, and more☆2,374Aug 19, 2026Updated 2 weeks ago
- Design a Low-cost-AI-Accelerator based on Google's Tensor Processing Unit Version 1.☆23Aug 24, 2026Updated last week
- ☆20Jan 13, 2026Updated 7 months ago
- An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.☆144Mar 21, 2026Updated 5 months ago