Full Transformer into a custom chip. microGPT in RTL, generating names on a Virtex-5 FPGA at ~56k tokens/second.
☆651Jun 25, 2026Updated 3 months ago
Alternatives and similar repositories for gateGPT
Users that are interested in gateGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- hardware implementation of transformers running microgpt at 50k+ tkps☆775May 14, 2026Updated 4 months ago
- Pipeline-parallel LLM inference across GPUs on separate machines.☆458Aug 9, 2026Updated last month
- A self-hosted, zero-knowledge, secure PHP notebook. Your notes, in your browser, on your server - every byte encrypted with AES-256 using…☆26Jun 20, 2026Updated 3 months ago
- [DATE'2025, TCAD'2025] Terafly : A Multi-Node FPGA Based Accelerator Design for Efficient Cooperative Inference in LLMs☆40Nov 13, 2025Updated 10 months ago
- hardware accelerator for deep convolutional neural networks☆74Feb 25, 2026Updated 7 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- microGPT benchmarks: a single M4 Max MacBook Pro P-core in C runs Karpathy's 4192-parameter transformer at ~71x the throughput of TALOS-V…☆164May 4, 2026Updated 5 months ago
- HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.☆2,137Sep 21, 2026Updated 2 weeks ago
- Training neural networks on Apple Neural Engine via reverse-engineered private APIs☆7,260Mar 10, 2026Updated 6 months ago
- Tmux sidebar for vibe coding. Manage sessions and monitor agents at a glance☆15Sep 13, 2026Updated 3 weeks ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,901Updated this week
- A cheap multipurpose mobile robotic platform for research, development, education and fun. All the code, data, instructions and files yo…☆16Jul 1, 2026Updated 3 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆23,698Updated this week
- Secure encrypted one-time secret sharing for clearnet and Tor onion services☆33Jun 15, 2026Updated 3 months ago
- PCCX v002 KV260 board/model integration and verification, consuming a pinned reusable core. Operated by Altifigence.☆20Sep 27, 2026Updated last week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- RISC-V Superscalar Educational Simulator based on Tomasulo's Algorithm☆36Nov 1, 2025Updated 11 months ago
- ☆29Aug 25, 2026Updated last month
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,143Aug 18, 2026Updated last month
- A Top-Down Profiler for GPU Applications☆24Feb 29, 2024Updated 2 years ago
- Using Feature Decomposition method to accelerate GNN inference☆13Sep 27, 2021Updated 5 years ago
- FSA: Fusing FlashAttention within a Single Systolic Array☆203Apr 15, 2026Updated 5 months ago
- OpenAI's privacy filter NER model architecture implemented in a minimal C++/GGML runtime☆303Jul 2, 2026Updated 3 months ago
- MAVLink Military Messages☆24Feb 19, 2026Updated 7 months ago
- LoRa chip into a coherent linear-FM chirp generator☆110Apr 24, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- reverse engineering the best-selling drones on Amazon to control programmatically☆572Jun 23, 2026Updated 3 months ago
- Open-source AI Accelerator Stack integrating compute, memory, and software — from RTL to PyTorch.☆29Updated this week
- Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B☆1,581Aug 14, 2026Updated last month
- Native insanely fast terminal built for agents☆75Sep 11, 2026Updated 3 weeks ago
- Ventus GPGPU ISA Simulator Based on Spike☆56Sep 23, 2026Updated 2 weeks ago
- PCIe stack written in Amaranth.☆18Apr 27, 2026Updated 5 months ago
- A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1☆1,392Apr 3, 2026Updated 6 months ago
- ☆175Updated this week
- A vector index built on TurboQuant, written in Rust with Python bindings☆17,358Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Minimal TPU implementation with 8x8 systolic array and PyTorch integration☆71Jan 26, 2026Updated 8 months ago
- Fast and memory-efficient classical machine learning operators☆596Aug 31, 2026Updated last month
- A private, open-source alternative to Windows Recall — for Linux. Captures your screen, OCRs it, and makes everything you've seen instant…☆55Sep 22, 2026Updated 2 weeks ago
- On-device semantic search over Apple WWDC 2025 docs using MLX embeddings — SwiftUI app (WWDC OMT 2025)☆76Jun 12, 2025Updated last year
- Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦☆40,573Updated this week
- ☆26Jul 16, 2026Updated 2 months ago
- ActiveGraph/GBrain bridge proof of concept for Apprentice launch.☆25May 26, 2026Updated 4 months ago