implement GPT-OSS 20B & 120B C++ inference from scratch on AMD GPUs
☆174Oct 25, 2025Updated 9 months ago
Alternatives and similar repositories for gpt-oss-amd
Users that are interested in gpt-oss-amd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- v0.1.0-alpha☆22Dec 2, 2025Updated 7 months ago
- Efficient Class Incremental Learning for Object Detection☆32Jun 30, 2024Updated 2 years ago
- TP - AI Project S2T1, DSAI HUST☆19Jun 7, 2022Updated 4 years ago
- DSLab Training☆27Jul 15, 2024Updated 2 years ago
- Machine Learning Experience☆21May 6, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A curated learning repository focused on High-Performance Computing (HPC) — covering fundamentals to advanced topics in CUDA, MPI, C++, a…☆363Mar 22, 2026Updated 4 months ago
- Load and run Llama from safetensors files in C☆15Oct 24, 2024Updated last year
- Developing K - a language model to generate OPENSCAD code from prompt☆19Dec 3, 2025Updated 7 months ago
- Protocol for Augmented Memory of Project Artifacts (MCP compatible) - extended☆24Jan 24, 2026Updated 6 months ago
- ☆15Feb 23, 2025Updated last year
- ☆27Jan 22, 2026Updated 6 months ago
- 3.34× faster inference on Apple Silicon — native MLX port of DFlash speculative decoding☆19Apr 11, 2026Updated 3 months ago
- Vibe 🎶 Coding 💻 Prolog 🐪☆56Jan 15, 2026Updated 6 months ago
- Composition of Multimodal Language Models From Scratch☆15Aug 16, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- amdgpu example code in hip/asm☆66Updated this week
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- ojjson is a library designed to facilitate JSON interactions with Ollama, a large language api (LLM). It leverages the power of Zod for s…☆12Nov 7, 2024Updated last year
- Optiflop measures the optimally achievable FLOPs for mathematical operations on various platforms.☆14Nov 1, 2024Updated last year
- Production-ready expert load balancing library☆80Jul 17, 2026Updated last week
- ☆16Jul 21, 2026Updated last week
- Scripts for building libraries with Cray's PE☆21Aug 31, 2021Updated 4 years ago
- Official implementation of "From Implicit to Explicit Feedback: A deep neural network for modeling sequential behaviours and long-short t…☆19Oct 16, 2025Updated 9 months ago
- From baby GPT to diffusion GPT: An annotated implementation of a character-level discrete diffusion model (adapted from Karpathy’s baby G…☆259Oct 12, 2025Updated 9 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ExpertFingerprinting: Behavioral Pattern Analysis and Specialization Mapping of Experts in GPT-OSS-20B's Mixture-of-Experts Architecture☆27Feb 3, 2026Updated 5 months ago
- MOL — The cognitive programming language with auto-tracing pipelines. Built for AI/RAG by CruxLabx.☆73Apr 11, 2026Updated 3 months ago
- Ichigo Whisper is a compact (22M parameters), open-source speech tokenizer for the Whisper-medium, designed to enhance performance on mul…☆16Jan 20, 2025Updated last year
- nvloom is a set of tools designed to scalably test MNNVL fabrics.☆50Updated this week
- Collection of memory microbenchmarks to investigate NVIDIA GPUs Network on Chip architectures☆15Apr 14, 2026Updated 3 months ago
- ☆27Aug 16, 2025Updated 11 months ago
- Egg Framework Boilerplate for NodeJS Runtime.☆14Mar 16, 2022Updated 4 years ago
- Convert the Berkeley Deepdrive dataset to a TFRecord file☆16May 8, 2019Updated 7 years ago
- A MCP stdio toolpack for local LLMs☆33Apr 6, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆58Aug 26, 2024Updated last year
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆18Feb 4, 2026Updated 5 months ago
- High-Performance FP32 GEMM on CUDA devices☆126Jan 21, 2025Updated last year
- Evaluation of an SGS model for mass transfer at risining bubbles in the initial transient stage☆13Mar 18, 2022Updated 4 years ago
- A repository of CrayLabs and user contributed examples of using SmartSim.☆18Jun 11, 2024Updated 2 years ago
- Note about running ollama 🦙☆36May 2, 2024Updated 2 years ago
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated 3 weeks ago