Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
☆484Jul 24, 2026Updated this week
Alternatives and similar repositories for krasis
Users that are interested in krasis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆18May 6, 2026Updated 2 months ago
- Terminal system resource monitor for hybrid LLM workloads☆83May 6, 2026Updated 2 months ago
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆18Feb 4, 2026Updated 5 months ago
- Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placem…☆257Updated this week
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆145Jul 14, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- llama.cpp fork with additional SOTA quants and improved performance☆2,960Updated this week
- Coding agent with cool features.☆16Jun 6, 2026Updated last month
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,129Updated this week
- Neural-enhanced conversational memory for AI agents — 10 micro-networks, tri-hybrid storage, bio-inspired retrieval☆15Mar 27, 2026Updated 3 months ago
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆77Jul 12, 2026Updated last week
- Zero-shot forecasting, tabular classification, and regression via MCP — exposes Google TimesFM 2.5 and TabFM v1.0.0 to AI assistants. Jus…☆24Jul 12, 2026Updated last week
- Fast LLM speculative inference server for consumer hardware.☆2,673Updated this week
- A tiny model that teaches itself to code better. On your laptop. No cloud. No teacher model. No human feedback.☆65Mar 10, 2026Updated 4 months ago
- A local-first web search agent☆29Jun 20, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 8 months ago
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated 2 weeks ago
- Using LLMs for iteratively exploring the solution search space at scale.☆734Jul 13, 2026Updated last week
- ☆83Feb 28, 2025Updated last year
- A collection of high quality huggingface datasets.☆29Apr 19, 2026Updated 3 months ago
- ☆19Jul 4, 2025Updated last year
- ☆1,566Updated this week
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆17Jul 3, 2026Updated 3 weeks ago
- An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on yo…☆2,270Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,778Updated this week
- A harness optimized to smaller LLMs☆1,877Updated this week
- Autonomous multi-agent coding system on local LLMs. No frameworks, no API costs. Plan → Build → Test → Fix with Qwen 80B + 7B on your own…☆20Feb 18, 2026Updated 5 months ago
- vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, …☆16Apr 26, 2026Updated 2 months ago
- Zero-instrumentation LLM API and MCP tracer for your agents powered by eBPF — latency, tokens, and tool use in realtime☆18Mar 16, 2026Updated 4 months ago
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated last month
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆146Updated this week
- autoresearch for everything — autonomous iterative improvement for any system☆37Apr 2, 2026Updated 3 months ago
- A TypeScript-based MCP-server tool enabling concurrent chains of thought with real-time reinforcement learning. Seamlessly integrates wit…☆21Mar 17, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Local-first AI memory with O(k) prefix queries. Your data, your machine.☆91Mar 23, 2026Updated 4 months ago
- Advanced interoperability middleware for GPGPU acceleration. Facilitates cross vendor hardware abstraction and API translation for parall…☆97May 23, 2026Updated 2 months ago
- High-performance LLM inference engine — drop-in replacement for Ollama with faster multi-turn inference, lower TTFT, and higher throughpu…☆169Updated this week
- Fully local, persona-driven AI companion bot for Telegram (Telethon) with optional Discord bridge. Modular text pipeline, per-user memory…☆17Jun 21, 2026Updated last month
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆165Updated this week
- Fast, lossless LLM inference via dual-view diffusion decoding.☆460May 18, 2026Updated 2 months ago
- AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.☆1,998Updated this week