High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.
☆18Mar 20, 2026Updated 5 months ago
Alternatives and similar repositories for fast_topk_batched
Users that are interested in fast_topk_batched are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A conversational AI system using Ollama with persistent memory capabilities. Features hybrid context management (sliding window + vector …☆21Mar 20, 2026Updated 5 months ago
- An Efficent BPE Algorithm Faster then Hugging Face Tokenizer's Implementation☆13Sep 9, 2024Updated last year
- Natural language control for Python CLI tools using locally-trained SLMs (CPU inference)☆33Aug 15, 2026Updated 2 weeks ago
- An example repo full of Douglas Adams quotes☆16Mar 12, 2024Updated 2 years ago
- Learn anything with Local Personalized Learning AI App☆29Aug 4, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- finetune method to create think/model/requires tags to allow LLMs to write programs for things they can calculate instead of hallucinatin…☆15Apr 8, 2026Updated 4 months ago
- The UnisonAI Multi-Agent Framework built on custom workflow which allows ai agents to talk together and provides a flexible and extensibl…☆23Feb 24, 2026Updated 6 months ago
- A tiny application that disables the Windows keys of your keyboard. Very useful in games!☆28May 12, 2023Updated 3 years ago
- Sherpa-onnx-tts-stt source for homeassisstant addon with Kroko Onnx Streaming STT integration.☆31Dec 18, 2025Updated 8 months ago
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 6 months ago
- CarbonMU was experiments towards a general-purpose, extendable MUD/MUSH server written in Ruby.☆13Aug 26, 2017Updated 9 years ago
- ☆11Sep 13, 2022Updated 3 years ago
- An old-school computer roleplaying game inspired by the early Ultima games written with LÖVE; The tileset used was created by Josh Steele…☆10Apr 30, 2018Updated 8 years ago
- An emulator of x86 assembly with indicators for registers and memories on web browser.☆11Dec 15, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Kolosal AI is an OpenSource and Lightweight alternative to Ollama to run LLMs 100% offline on your device.☆15Jan 2, 2026Updated 7 months ago
- GPU-accelerated voice assistant — local LLM, fine-tuned Whisper, Kokoro TTS, AMD ROCm☆25Apr 2, 2026Updated 4 months ago
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆22Feb 16, 2026Updated 6 months ago
- Automated parameter sweep pipeline for finding optimal sampling settings for any local LLM on quantized weights☆22Feb 21, 2026Updated 6 months ago
- Offline LLM chatbot with personalized memory — works on CPU with multi-session memory support.☆22Jan 10, 2026Updated 7 months ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- So, I trained a Llama a 130M architecture I coded from ground up to build a small instruct model from scratch. Trained on FineWeb dataset…☆18Mar 26, 2025Updated last year
- Real time faster whisper gradio☆24Aug 17, 2025Updated last year
- This is creating a TCP Server using Rust programming language☆14Oct 3, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆18Apr 19, 2024Updated 2 years ago
- ☆12Jan 22, 2016Updated 10 years ago
- ☆11Mar 18, 2026Updated 5 months ago
- Deterministic-mode checks for LLM inference: measure run/batch variance, generate repro packs, and explain why outputs differ.☆20Aug 20, 2026Updated last week
- A Dwarf Fortress clone☆18Feb 15, 2018Updated 8 years ago
- JAX port of FLUX.1 models using flax.nnx☆23Sep 28, 2024Updated last year
- ☆13Jul 9, 2024Updated 2 years ago
- SLMs for personal expenses summaries☆23Dec 20, 2025Updated 8 months ago
- Sample game project using the Godot engine.☆14Feb 27, 2019Updated 7 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Huge f*cking file delta copy☆10Aug 3, 2025Updated last year
- The exercise for the user level thread programming☆14Dec 15, 2020Updated 5 years ago
- Plug-and-play terminal security layer for LLM agents. Drop-in gatekeeper that prevents dangerous shell commands. Works with OpenAI, Claud…☆24Jan 29, 2026Updated 7 months ago
- Pressure-test your specs with LLM reasoning before writing code. Agent skill for Claude Code, Codex, Gemini CLI, and 14+ coding agents.☆22Mar 22, 2026Updated 5 months ago
- An API for VoiceCraft.☆25Jun 27, 2024Updated 2 years ago
- A fork of the GroundZero2 codebase with bug fixes, standardization, and sanity-related changes.☆11Jan 7, 2022Updated 4 years ago
- @ swap drop (forth rogue). docforth processed code at☆16Jan 1, 2019Updated 7 years ago