eLLM can infer LLM on CPUs faster than on GPUs
☆438Aug 21, 2026Updated this week
Alternatives and similar repositories for eLLM
Users that are interested in eLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Speculative Decoding Implementations: MTP, EAGLE-3, Medusa-1, PARD, Draft Models, N-gram and Suffix Decoding from scratch☆15May 2, 2026Updated 3 months ago
- A 15M parameter character LLM trained to survive in an empty world with only a pigeon for company.☆22May 5, 2026Updated 3 months ago
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- ☆2,784Jul 27, 2026Updated 3 weeks ago
- A from-scratch LLM inference engine and chat application.☆47May 23, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,898Updated this week
- Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free…☆25Aug 4, 2026Updated 2 weeks ago
- Self-hosted, OpenAI-compatible RAG API + MCP server that plugs local knowledge into existing LLM clients.☆33Updated this week
- Adaptive Chunking: automatically select the best chunking method per document for RAG. Accepted at LREC 2026.☆376Jul 6, 2026Updated last month
- GLiNER2 Rust support☆20Aug 16, 2026Updated last week
- Bias, Hate classification with KoELECTRA 👿☆27Jun 12, 2023Updated 3 years ago
- EdegQuake 🌋 High-performance GraphRAG inspired from LightRag written in Rust; Transform documents into intelligent knowledge graphs for …☆2,074Updated this week
- ☆20Jan 3, 2026Updated 7 months ago
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆204May 15, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Row-Bot - Personal AI Sovereignty. A local-first AI assistant with integrated tools, a personal knowledge graph, voice, vision, shell, br…☆1,444Updated this week
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆24May 5, 2026Updated 3 months ago
- Can AI Agents Build Bespoke Systems?☆92Updated this week
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch…☆134Jun 29, 2026Updated last month
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,322Updated this week
- the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for mi…☆121Jul 7, 2026Updated last month
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,227Jul 30, 2026Updated 3 weeks ago
- Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.☆74,496Updated this week
- Your ai-powered everything assistant. Oboto is an AI assistant that runs on your computer and in the cloud. Oboto remembers who you are a…☆27May 8, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 4 months ago
- Creation OS - Cognitive Architecture☆42Jun 24, 2026Updated 2 months ago
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆149Jul 27, 2026Updated 3 weeks ago
- Architecture pattern for combining a fast LLM voice loop with a slower SLM that tracks hard facts.☆15Apr 27, 2026Updated 3 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,532Mar 19, 2026Updated 5 months ago
- A cognitive architecture that runs on your own machine. Internal state reaches generation through the model's activations, not the system…☆76Updated this week
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆1,466Updated this week
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆280Updated this week
- A transparent (O)llama proxy with model deployment aware routing which auto-manages multiple (O)llama instances in a given network.☆20Apr 14, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 4 months ago
- ☆12Feb 17, 2025Updated last year
- ☆23May 30, 2025Updated last year
- repo-grep performs a recursive grep through the folder structure of your cloned Git repository or your SVN working copy in Emacs. By defa…☆21Jun 6, 2026Updated 2 months ago
- Stable Looped Models and their Scaling Laws☆174May 17, 2026Updated 3 months ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,040Apr 23, 2026Updated 4 months ago
- SuperOptiX: Full Stack Agentic AI Framework☆24May 2, 2026Updated 3 months ago