Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
☆521Sep 15, 2026Updated last week
Alternatives and similar repositories for krasis
Users that are interested in krasis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆21Feb 4, 2026Updated 7 months ago
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆275Updated this week
- Autonomous multi-agent coding system on local LLMs. No frameworks, no API costs. Plan → Build → Test → Fix with Qwen 80B + 7B on your own…☆21Updated this week
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆157Sep 8, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- llama.cpp fork with additional SOTA quants and improved performance☆3,254Updated this week
- Coding agent with cool features.☆19Aug 23, 2026Updated 3 weeks ago
- An MCP server to read MCP logs to debug directly inside the client☆16May 3, 2025Updated last year
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,722Updated this week
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆85Jul 12, 2026Updated 2 months ago
- Zero-shot forecasting, tabular classification, and regression via MCP — exposes Google TimesFM 2.5 and TabFM v1.0.0 to AI assistants. Jus…☆29Sep 4, 2026Updated 2 weeks ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,872Updated this week
- A tiny model that teaches itself to code better. On your laptop. No cloud. No teacher model. No human feedback.☆70Mar 10, 2026Updated 6 months ago
- A local-first web search agent☆30Jun 20, 2026Updated 3 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated 2 months ago
- Using LLMs for iteratively exploring the solution search space at scale.☆739Sep 5, 2026Updated 2 weeks ago
- ☆84Feb 28, 2025Updated last year
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 10 months ago
- A collection of high quality huggingface datasets.☆30Apr 19, 2026Updated 5 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,280Updated this week
- Why observe computer if computer can observe for you☆1,620Updated this week
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆18Sep 6, 2026Updated 2 weeks ago
- An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on yo…☆2,489Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A harness optimized to smaller LLMs☆2,611Updated this week
- Zero-instrumentation LLM API and MCP tracer for your agents powered by eBPF — latency, tokens, and tool use in realtime☆18Mar 16, 2026Updated 6 months ago
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆60Jun 10, 2026Updated 3 months ago
- An AI-assisted coding agent that runs in your terminal and supports many providers and models.☆231Updated this week
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆168Sep 11, 2026Updated last week
- autoresearch for everything — autonomous iterative improvement for any system☆38Apr 2, 2026Updated 5 months ago
- A TypeScript-based MCP-server tool enabling concurrent chains of thought with real-time reinforcement learning. Seamlessly integrates wit…☆21Mar 17, 2025Updated last year
- Local-first AI memory with O(k) prefix queries. Your data, your machine.☆90Mar 23, 2026Updated 5 months ago
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆18Apr 26, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fully local, persona-driven AI companion bot for Telegram (Telethon) with optional Discord bridge. Modular text pipeline, per-user memory…☆20Jun 21, 2026Updated 3 months ago
- Fast, lossless LLM inference via dual-view diffusion decoding.☆481Aug 12, 2026Updated last month
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆177Updated this week
- AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.☆2,027Aug 12, 2026Updated last month
- RDNA-native LLM inference engine in Rust.☆635Updated this week
- LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.☆402Apr 26, 2026Updated 4 months ago
- Your AI's anchor to reality. ⚓☆61Updated this week