eLLM: Run Long-Horizon Inference Faster on CPUs Than on GPUs
☆590Sep 10, 2026Updated this week
Alternatives and similar repositories for eLLM
Users that are interested in eLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A small Python library and CLI for working with Cloudflare Turnstile on your own pages - read the sitekey, build the token, verify the re…☆126Sep 6, 2026Updated last week
- A small Python library and CLI for working with Cloudflare Turnstile on your own pages - read the sitekey, build the token, verify the re…☆317Sep 6, 2026Updated last week
- A 15M parameter character LLM trained to survive in an empty world with only a pigeon for company.☆21May 5, 2026Updated 4 months ago
- ☆2,798Jul 27, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,852Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Free and Open Source version of Scale Space, a powerful, feature-rich cockpit to explore the expanse of phase space.☆20May 28, 2026Updated 3 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,083Aug 18, 2026Updated 3 weeks ago
- rotating proxy system☆25Updated this week
- A from-scratch LLM inference engine and chat application.☆47May 23, 2026Updated 3 months ago
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch…☆139Jun 29, 2026Updated 2 months ago
- the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for mi…☆125Jul 7, 2026Updated 2 months ago
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆203May 15, 2026Updated 3 months ago
- Self-hosted, OpenAI-compatible RAG API + MCP server that plugs local knowledge into existing LLM clients.☆34Updated this week
- GLiNER2 Rust support☆21Aug 29, 2026Updated 2 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Adaptive Chunking: automatically select the best chunking method per document for RAG. Accepted at LREC 2026.☆390Jul 6, 2026Updated 2 months ago
- Row-Bot - Personal AI Sovereignty. A local-first AI assistant with integrated tools, a personal knowledge graph, voice, vision, shell, br…☆1,491Updated this week
- EdegQuake 🌋 High-performance GraphRAG inspired from LightRag written in Rust; Transform documents into intelligent knowledge graphs for …☆2,092Updated this week
- ☆20Jan 3, 2026Updated 8 months ago
- ☆14Aug 27, 2020Updated 6 years ago
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆87Jun 23, 2026Updated 2 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,802Updated this week
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.☆76,148Updated this week
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆283Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆26May 5, 2026Updated 4 months ago
- Can AI Agents Build Bespoke Systems?☆97Updated this week
- Perplexity style AI answer engine for AI PCs with CPU,GPU and NPU support☆51Mar 1, 2026Updated 6 months ago
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,250Sep 1, 2026Updated 2 weeks ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,555Mar 19, 2026Updated 5 months ago
- Your ai-powered everything assistant. Oboto is an AI assistant that runs on your computer and in the cloud. Oboto remembers who you are a…☆27May 8, 2026Updated 4 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,047Apr 23, 2026Updated 4 months ago
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆156Updated this week
- Architecture pattern for combining a fast LLM voice loop with a slower SLM that tracks hard facts.☆15Apr 27, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,046Apr 23, 2026Updated 4 months ago
- A vector index built on TurboQuant, written in Rust with Python bindings☆17,128Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆8,071Updated this week
- Resource monitor for Machine Learning tasks written in Rust☆14Sep 4, 2026Updated last week
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 5 months ago
- ☆12Feb 17, 2025Updated last year
- ☆23May 30, 2025Updated last year