eLLM can infer LLM on CPUs faster than on GPUs
☆428Jul 30, 2026Updated this week
Alternatives and similar repositories for eLLM
Users that are interested in eLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Speculative Decoding Implementations: MTP, EAGLE-3, Medusa-1, PARD, Draft Models, N-gram and Suffix Decoding from scratch☆15May 2, 2026Updated 3 months ago
- ☆32Jun 18, 2026Updated last month
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- Expose what functional RTL benchmarks leave unanswered. Evidence profiles for AI-generated RTL; research collaborators and design partner…☆18Jul 22, 2026Updated last week
- ☆2,764Jul 27, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Free and Open Source version of Scale Space, a powerful, feature-rich cockpit to explore the expanse of phase space.☆20May 28, 2026Updated 2 months ago
- rotating proxy system☆22Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,562May 10, 2026Updated 2 months ago
- Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free…☆24Jul 26, 2026Updated last week
- Self-hosted, OpenAI-compatible RAG API + MCP server that plugs local knowledge into existing LLM clients.☆33Updated this week
- Adaptive Chunking: automatically select the best chunking method per document for RAG. Accepted at LREC 2026.☆375Jul 6, 2026Updated 3 weeks ago
- ☆20Jan 3, 2026Updated 7 months ago
- ☆14Aug 27, 2020Updated 5 years ago
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆77Jun 23, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆200May 15, 2026Updated 2 months ago
- Row-Bot - Personal AI Sovereignty. A local-first AI assistant with integrated tools, a personal knowledge graph, voice, vision, shell, br…☆1,404Updated this week
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆23May 5, 2026Updated 2 months ago
- Can AI Agents Build Bespoke Systems?☆89Updated this week
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch…☆125Jun 29, 2026Updated last month
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,000Updated this week
- the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for mi…☆120Jul 7, 2026Updated 3 weeks ago
- Perplexity style AI answer engine for AI PCs with CPU,GPU and NPU support☆50Mar 1, 2026Updated 5 months ago
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,147Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.☆69,461Updated this week
- Your ai-powered everything assistant. Oboto is an AI assistant that runs on your computer and in the cloud. Oboto remembers who you are a…☆27May 8, 2026Updated 2 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 3 months ago
- Creation OS - Cognitive Architecture☆42Jun 24, 2026Updated last month
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆148Jul 27, 2026Updated last week
- Browser-native vector database built in Rust and compiled to WebAssembly for client-side embedding storage and semantic search.☆18Mar 13, 2026Updated 4 months ago
- Architecture pattern for combining a fast LLM voice loop with a slower SLM that tracks hard facts.☆15Apr 27, 2026Updated 3 months ago
- A cognitive architecture that runs on your own machine. Internal state reaches generation through the model's activations, not the system…☆74Updated this week
- A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do☆1,278Apr 27, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆280Jul 6, 2026Updated 3 weeks ago
- A transparent (O)llama proxy with model deployment aware routing which auto-manages multiple (O)llama instances in a given network.☆20Apr 14, 2026Updated 3 months ago
- Resource monitor for Machine Learning tasks written in Rust☆14Updated this week
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 3 months ago
- ☆12Feb 17, 2025Updated last year
- ☆18Feb 25, 2026Updated 5 months ago
- From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% a…☆1,038Apr 23, 2026Updated 3 months ago