Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
☆510Aug 13, 2026Updated this week
Alternatives and similar repositories for krasis
Users that are interested in krasis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- Terminal system resource monitor for hybrid LLM workloads☆84May 6, 2026Updated 3 months ago
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆21Feb 4, 2026Updated 6 months ago
- llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.☆264Updated this week
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆149Jul 27, 2026Updated 2 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- llama.cpp fork with additional SOTA quants and improved performance☆3,028Updated this week
- Coding agent with cool features.☆19Jun 6, 2026Updated 2 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,345Updated this week
- Neural-enhanced conversational memory for AI agents — 10 micro-networks, tri-hybrid storage, bio-inspired retrieval☆15Mar 27, 2026Updated 4 months ago
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆83Jul 12, 2026Updated last month
- Zero-shot forecasting, tabular classification, and regression via MCP — exposes Google TimesFM 2.5 and TabFM v1.0.0 to AI assistants. Jus…☆27Jul 12, 2026Updated last month
- LLM speculative inference server for consumer & heterogeneous hardware☆2,739Updated this week
- A tiny model that teaches itself to code better. On your laptop. No cloud. No teacher model. No human feedback.☆68Mar 10, 2026Updated 5 months ago
- A local-first web search agent☆29Jun 20, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 9 months ago
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated last month
- Using LLMs for iteratively exploring the solution search space at scale.☆744Updated this week
- ☆84Feb 28, 2025Updated last year
- A collection of high quality huggingface datasets.☆30Apr 19, 2026Updated 3 months ago
- ☆19Jul 4, 2025Updated last year
- ☆1,597Updated this week
- An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on yo…☆2,346Updated this week
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,932Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A harness optimized to smaller LLMs☆2,373Jul 31, 2026Updated last week
- Autonomous multi-agent coding system on local LLMs. No frameworks, no API costs. Plan → Build → Test → Fix with Qwen 80B + 7B on your own…☆20Feb 18, 2026Updated 5 months ago
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆16Apr 26, 2026Updated 3 months ago
- Zero-instrumentation LLM API and MCP tracer for your agents powered by eBPF — latency, tokens, and tool use in realtime☆18Mar 16, 2026Updated 4 months ago
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆60Jun 10, 2026Updated 2 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Aug 1, 2026Updated last week
- An AI-assisted coding agent that runs in your terminal and supports many providers and models.☆229Updated this week
- autoresearch for everything — autonomous iterative improvement for any system☆38Apr 2, 2026Updated 4 months ago
- Local-first AI memory with O(k) prefix queries. Your data, your machine.☆90Mar 23, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- High-performance LLM inference engine — drop-in replacement for Ollama with faster multi-turn inference, lower TTFT, and higher throughpu…☆186Updated this week
- Fully local, persona-driven AI companion bot for Telegram (Telethon) with optional Discord bridge. Modular text pipeline, per-user memory…☆18Jun 21, 2026Updated last month
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆168Updated this week
- Fast, lossless LLM inference via dual-view diffusion decoding.☆474Updated this week
- AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.☆2,006Updated this week
- RDNA-native LLM inference engine in Rust.☆511Updated this week
- A Windows tool to query various LLM AIs. Supports branched conversations, history and summaries among others.☆36May 11, 2026Updated 3 months ago