From-scratch implementation of OpenAI's GPT-OSS model in Python. No Torch, No GPUs.
☆110Nov 5, 2025Updated 8 months ago
Alternatives and similar repositories for GPT-OSS
Users that are interested in GPT-OSS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bloat Free, Portable and Lightweight LLM Frontend (Single HTML file). With Lorebook, Web Search, Macro Engine etc.☆22Jul 18, 2026Updated last week
- A Windows tool to query various LLM AIs. Supports branched conversations, history and summaries among others.☆36May 11, 2026Updated 2 months ago
- Random llm scripts☆41May 26, 2026Updated 2 months ago
- A lightweight chat interface for interacting with local models, featuring persistent memory using a seamless SQLite database to store you…☆34Sep 15, 2025Updated 10 months ago
- Local-first agent runtime for MCP workflows with explicit trust controls, replayable runs, and built-in evals.☆32Jul 4, 2026Updated 3 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- FastAPI + MLX offline-first voice agent with <1s latency. Minimal UI☆55Oct 21, 2025Updated 9 months ago
- Local-first RAG application for technical documentation and research papers☆28Jun 26, 2026Updated last month
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- ☆20Jan 3, 2026Updated 6 months ago
- Local modular AI assistant with speech, vision, and robotics support. Uses Qwen3-VL-4B in LM Studio.☆53Jan 9, 2026Updated 6 months ago
- Offline LLM chatbot with personalized memory — works on CPU with multi-session memory support.☆22Jan 10, 2026Updated 6 months ago
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 6 months ago
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆25Apr 1, 2025Updated last year
- ☆18Jul 1, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A natural language file search tool that uses LLMs to help you find files by describing what you're looking for.☆28Mar 8, 2025Updated last year
- Local AI chat application with automatic offline context injection through ZIM files.☆47Apr 3, 2026Updated 3 months ago
- 🤖 Google AI Mode Scraper - Get instant AI answers with beautiful table formatting. Perfect for students, researchers, and educators. Wor…☆22Nov 30, 2025Updated 7 months ago
- High-Performance Text Deduplication Toolkit☆61Aug 25, 2025Updated 11 months ago
- smallevals — CPU-fast, GPU-blazing fast offline retrieval evaluation for RAG systems with tiny QA models.☆22Dec 4, 2025Updated 7 months ago
- An fully autonomous agent that accesses the browser and performs tasks.☆18Apr 25, 2025Updated last year
- ☆20Jul 4, 2025Updated last year
- Local Qwen3 LLM inference. One easy-to-understand file of C source with no dependencies.☆184Jul 5, 2025Updated last year
- ☆15Mar 18, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- German Text Embedding Clustering Benchmark☆19Mar 15, 2024Updated 2 years ago
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆35Apr 20, 2026Updated 3 months ago
- Would you like to recreate GPT-2 124m in a cave with a box of scraps and a 4090 in less than two hours? LETS SPEEDRUN!☆39Nov 26, 2025Updated 8 months ago
- A truly open version of gpt-oss which shows the entire pre-training from scratch☆90Sep 4, 2025Updated 10 months ago
- An agent that can run everywhere - even in your watch!☆34Apr 8, 2026Updated 3 months ago
- Load and run Llama from safetensors files in C☆15Oct 24, 2024Updated last year
- CustomGPT.ai’s RAG API’s Starter Kit, including multi-instance embedded widgets, floating buttons, and standalone application.☆52Dec 23, 2025Updated 7 months ago
- CPU-Native Language Models☆32May 18, 2026Updated 2 months ago
- interactive semantic search demo using Qwen3-0.6B-Embedding in your browser☆60Feb 25, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Python console application designed to provide an engaging and visually appealing LLM chat experience on Unix-like consoles or Terminals.☆26May 20, 2026Updated 2 months ago
- Super simple python connectors for llama.cpp, including vision models (Gemma 3, Qwen2-VL). Compile llama.cpp and run!☆31Dec 11, 2025Updated 7 months ago
- SwiftLet is a lightweight Python framework for running open-source Large Language Models (LLMs) locally using safetensors☆29Aug 6, 2025Updated 11 months ago
- Constitutional Physics Framework for Digital Consciousness☆18Nov 16, 2025Updated 8 months ago
- This Streamlit application allows users to upload images and engage in interactive conversations about them using the Ollama Vision Model…☆15Nov 11, 2024Updated last year
- Code for the paper: "StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation"☆40May 16, 2025Updated last year
- Production-ready ternary quantized (1.58-bit) Rust code generation model with mHC-lite, MaxRL training, and comprehensive benchmarking☆21Mar 7, 2026Updated 4 months ago