High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.
☆17Mar 20, 2026Updated 4 months ago
Alternatives and similar repositories for fast_topk_batched
Users that are interested in fast_topk_batched are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A conversational AI system using Ollama with persistent memory capabilities. Features hybrid context management (sliding window + vector …☆18Mar 20, 2026Updated 4 months ago
- An Efficent BPE Algorithm Faster then Hugging Face Tokenizer's Implementation☆13Sep 9, 2024Updated last year
- MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation☆22Feb 10, 2026Updated 5 months ago
- Natural language control for Python CLI tools using locally-trained SLMs (CPU inference)☆32Apr 10, 2026Updated 3 months ago
- ☆33May 15, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LlamaNet: Decentralized Inference Swarm for llama.cpp☆23Jun 19, 2026Updated last month
- Learn anything with Local Personalized Learning AI App☆28Mar 26, 2026Updated 3 months ago
- finetune method to create think/model/requires tags to allow LLMs to write programs for things they can calculate instead of hallucinatin…☆15Apr 8, 2026Updated 3 months ago
- The UnisonAI Multi-Agent Framework built on custom workflow which allows ai agents to talk together and provides a flexible and extensibl…☆23Feb 24, 2026Updated 4 months ago
- A tiny application that disables the Windows keys of your keyboard. Very useful in games!☆27May 12, 2023Updated 3 years ago
- collection of utilities for file manipulation in golang☆13Oct 4, 2023Updated 2 years ago
- Sherpa-onnx-tts-stt source for homeassisstant addon with Kroko Onnx Streaming STT integration.☆30Dec 18, 2025Updated 7 months ago
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 5 months ago
- ☆15Jan 6, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- CarbonMU was experiments towards a general-purpose, extendable MUD/MUSH server written in Ruby.☆13Aug 26, 2017Updated 8 years ago
- run ai without internet on web and mobile☆15Updated this week
- ☆11Sep 15, 2015Updated 10 years ago
- An old-school computer roleplaying game inspired by the early Ultima games written with LÖVE; The tileset used was created by Josh Steele…☆10Apr 30, 2018Updated 8 years ago
- An emulator of x86 assembly with indicators for registers and memories on web browser.☆11Dec 15, 2023Updated 2 years ago
- GPU-accelerated voice assistant — local LLM, fine-tuned Whisper, Kokoro TTS, AMD ROCm☆16Apr 2, 2026Updated 3 months ago
- Kolosal AI is an OpenSource and Lightweight alternative to Ollama to run LLMs 100% offline on your device.☆15Jan 2, 2026Updated 6 months ago
- Based on the AberMUD V source code by Alan Cox☆13Nov 21, 2016Updated 9 years ago
- Offline LLM chatbot with personalized memory — works on CPU with multi-session memory support.☆22Jan 10, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 2 months ago
- Real time faster whisper gradio☆24Aug 17, 2025Updated 11 months ago
- So, I trained a Llama a 130M architecture I coded from ground up to build a small instruct model from scratch. Trained on FineWeb dataset…☆18Mar 26, 2025Updated last year
- Deterministic-mode checks for LLM inference: measure run/batch variance, generate repro packs, and explain why outputs differ.☆19May 29, 2026Updated last month
- A Dwarf Fortress clone☆17Feb 15, 2018Updated 8 years ago
- SLMs for personal expenses summaries☆23Dec 20, 2025Updated 7 months ago
- Plug-and-play terminal security layer for LLM agents. Drop-in gatekeeper that prevents dangerous shell commands. Works with OpenAI, Claud…☆24Jan 29, 2026Updated 5 months ago
- Sample game project using the Godot engine.☆14Feb 27, 2019Updated 7 years ago
- Pollinations.ai Native Plugin for OpenCode. Access Free and Enterprise AI models directly in your editor.☆17Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Pressure-test your specs with LLM reasoning before writing code. Agent skill for Claude Code, Codex, Gemini CLI, and 14+ coding agents.☆21Mar 22, 2026Updated 3 months ago
- A miniaturized version of the Kimi-K2 model optimized for deployment on single H100 GPUs.☆35Jul 16, 2025Updated last year
- An API for VoiceCraft.☆25Jun 27, 2024Updated 2 years ago
- @ swap drop (forth rogue). docforth processed code at☆16Jan 1, 2019Updated 7 years ago
- Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).☆19Jun 9, 2026Updated last month
- A framework for creating multiplayer, text-based worlds.☆13Nov 15, 2019Updated 6 years ago
- A zero-config OpenAI client with support for 20+ providers, API key rotation, rate limits, optional LangChain integration and more.☆19Dec 11, 2025Updated 7 months ago