Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
☆516Sep 2, 2026Updated this week
Alternatives and similar repositories for krasis
Users that are interested in krasis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆21Feb 4, 2026Updated 6 months ago
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆268Updated this week
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆151Jul 27, 2026Updated last month
- llama.cpp fork with additional SOTA quants and improved performance☆3,171Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Coding agent with cool features.☆19Aug 23, 2026Updated last week
- An MCP server to read MCP logs to debug directly inside the client☆16May 3, 2025Updated last year
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,553Updated this week
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆84Jul 12, 2026Updated last month
- A low-cost and high-performance evolutionary engine for code and prompt optimization☆43Jun 22, 2026Updated 2 months ago
- Zero-shot forecasting, tabular classification, and regression via MCP — exposes Google TimesFM 2.5 and TabFM v1.0.0 to AI assistants. Jus…☆29Jul 12, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,828Updated this week
- A tiny model that teaches itself to code better. On your laptop. No cloud. No teacher model. No human feedback.☆68Mar 10, 2026Updated 5 months ago
- A local-first web search agent☆30Jun 20, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 9 months ago
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated last month
- Using LLMs for iteratively exploring the solution search space at scale.☆743Aug 15, 2026Updated 2 weeks ago
- ☆84Feb 28, 2025Updated last year
- A collection of high quality huggingface datasets.☆30Apr 19, 2026Updated 4 months ago
- ☆19Jul 4, 2025Updated last year
- ☆1,605Updated this week
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆18Jul 3, 2026Updated 2 months ago
- An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on yo…☆2,432Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,166Updated this week
- A harness optimized to smaller LLMs☆2,528Updated this week
- Zero-instrumentation LLM API and MCP tracer for your agents powered by eBPF — latency, tokens, and tool use in realtime☆18Mar 16, 2026Updated 5 months ago
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated 2 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆155Aug 21, 2026Updated last week
- autoresearch for everything — autonomous iterative improvement for any system☆38Apr 2, 2026Updated 5 months ago
- A TypeScript-based MCP-server tool enabling concurrent chains of thought with real-time reinforcement learning. Seamlessly integrates wit…☆21Mar 17, 2025Updated last year
- Local-first AI memory with O(k) prefix queries. Your data, your machine.☆90Mar 23, 2026Updated 5 months ago
- A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that sha…☆188Aug 22, 2026Updated last week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Fully local, persona-driven AI companion bot for Telegram (Telethon) with optional Discord bridge. Modular text pipeline, per-user memory…☆20Jun 21, 2026Updated 2 months ago
- Fast, lossless LLM inference via dual-view diffusion decoding.☆480Aug 12, 2026Updated 3 weeks ago
- AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.☆2,023Aug 12, 2026Updated 3 weeks ago
- RDNA-native LLM inference engine in Rust.☆591Updated this week
- A Windows tool to query various LLM AIs. Supports branched conversations, history and summaries among others.☆36May 11, 2026Updated 3 months ago
- A fully local, zero-API, zero-finetune multi-agent AI architecture that makes an 8B base model perform high level model reasoning, resear…☆34Nov 25, 2025Updated 9 months ago
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆399Apr 26, 2026Updated 4 months ago