Full Transformer into a custom chip. microGPT in RTL, generating names on a Virtex-5 FPGA at ~56k tokens/second.
☆645Jun 25, 2026Updated 2 months ago
Alternatives and similar repositories for gateGPT
Users that are interested in gateGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- hardware implementation of transformers running microgpt at 50k+ tkps☆776May 14, 2026Updated 4 months ago
- Pipeline-parallel LLM inference across GPUs on separate machines.☆456Aug 9, 2026Updated last month
- A self-hosted, zero-knowledge, secure PHP notebook. Your notes, in your browser, on your server - every byte encrypted with AES-256 using…☆26Jun 20, 2026Updated 2 months ago
- [DATE'2025, TCAD'2025] Terafly : A Multi-Node FPGA Based Accelerator Design for Efficient Cooperative Inference in LLMs☆39Nov 13, 2025Updated 10 months ago
- hardware accelerator for deep convolutional neural networks☆75Feb 25, 2026Updated 6 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- microGPT benchmarks: a single M4 Max MacBook Pro P-core in C runs Karpathy's 4192-parameter transformer at ~71x the throughput of TALOS-V…☆164May 4, 2026Updated 4 months ago
- Scout — the entire internet in one governed AI agent. Reads Twitter, Reddit, YouTube, GitHub & more via your local browser sessions. Read…☆26Jul 4, 2026Updated 2 months ago
- HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.☆2,045Sep 4, 2026Updated 2 weeks ago
- Training neural networks on Apple Neural Engine via reverse-engineered private APIs☆7,261Mar 10, 2026Updated 6 months ago
- Tmux sidebar for vibe coding. Manage sessions and monitor agents at a glance☆15Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,867Updated this week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆22,504Updated this week
- Secure encrypted one-time secret sharing for clearnet and Tor onion services☆33Jun 15, 2026Updated 3 months ago
- KV260 integration lane for PCCX™ v002 LLM IP-core bring-up, validation, and board/runtime evidence.☆19Jun 3, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- RISC-V Superscalar Educational Simulator based on Tomasulo's Algorithm☆36Nov 1, 2025Updated 10 months ago
- ☆27Aug 25, 2026Updated 3 weeks ago
- A collection of examples showcasing PyCDE and Mini RISC-V implementation.☆10Jul 23, 2026Updated last month
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,100Aug 18, 2026Updated last month
- A Top-Down Profiler for GPU Applications☆24Feb 29, 2024Updated 2 years ago
- FSA: Fusing FlashAttention within a Single Systolic Array☆198Apr 15, 2026Updated 5 months ago
- Free open-source extractor for AI coding assistant chat histories. Supports Claude Code, Cursor, Windsurf, Aider, Cline/Roo Code, and mor…☆840Sep 11, 2026Updated last week
- LoRa chip into a coherent linear-FM chirp generator☆108Apr 24, 2026Updated 4 months ago
- reverse engineering the best-selling drones on Amazon to control programmatically☆568Jun 23, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B☆1,576Aug 14, 2026Updated last month
- Native insanely fast terminal built for agents☆75Sep 11, 2026Updated last week
- Ventus GPGPU ISA Simulator Based on Spike☆53Sep 3, 2026Updated 2 weeks ago
- PCIe stack written in Amaranth.☆18Apr 27, 2026Updated 4 months ago
- ☆138Updated this week
- A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1☆1,375Apr 3, 2026Updated 5 months ago
- A vector index built on TurboQuant, written in Rust with Python bindings☆17,197Updated this week
- Minimal TPU implementation with 8x8 systolic array and PyTorch integration☆69Jan 26, 2026Updated 7 months ago
- Fast and memory-efficient classical machine learning operators☆592Aug 31, 2026Updated 2 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦☆36,062Updated this week
- On-device semantic search over Apple WWDC 2025 docs using MLX embeddings — SwiftUI app (WWDC OMT 2025)☆76Jun 12, 2025Updated last year
- ☆26Jul 16, 2026Updated 2 months ago
- ActiveGraph/GBrain bridge proof of concept for Apprentice launch.☆25May 26, 2026Updated 3 months ago
- Using Gemini Nano through Chrome's built-in Prompt API, with scripts to automate setup and verification☆179Jun 23, 2026Updated 2 months ago
- MathCode: A Frontier Mathematical Coding Agent☆741Sep 9, 2026Updated last week
- TokenSpeed is a speed-of-light LLM inference engine.☆2,146Updated this week