Pipeline-parallel LLM inference across GPUs on separate machines.
☆435Aug 3, 2026Updated this week
Alternatives and similar repositories for shard
Users that are interested in shard are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Uncensored, private, decentralized AI inference on Solana.☆63Updated this week
- Full Transformer into a custom chip. microGPT in RTL, generating names on a Virtex-5 FPGA at ~56k tokens/second.☆624Jun 25, 2026Updated last month
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- Distil SN97 — Competitive Model Distillation on Bittensor☆35May 20, 2026Updated 2 months ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆42Jun 28, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This repositories contains the reference implementation for the Sparse Delta Memory paper.More precisely, it contains the model definitio…☆32Jul 9, 2026Updated 3 weeks ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,868Updated this week
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆491Updated this week
- Mixed-vendor GPU inference cluster manager with speculative decoding☆31Jul 2, 2026Updated last month
- A private, open-source alternative to Windows Recall — for Linux. Captures your screen, OCRs it, and makes everything you've seen instant…☆45Jun 24, 2026Updated last month
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,526Updated this week
- An arbitrage bot is a smart contract connected to an external automation script that controls its operation.☆1,969Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Updated this week
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆18Jul 5, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside emb…☆37Updated this week
- Implementation of SelfExtend from the paper "LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning" from Pytorch and Zeta☆13Nov 11, 2024Updated last year
- Bonsai Demo☆2,149Updated this week
- Production-grade DSPy 3.2.x agent skills + validated end-to-end examples for Claude Code and Codex CLI — fundamentals, evaluation, GEPA, …☆267Jun 20, 2026Updated last month
- Single-file installer for a club-3090 webserver providing an admin control panel, a reverse proxy that automatically routes requests to t…☆22Jul 8, 2026Updated 3 weeks ago
- ☆19Mar 21, 2026Updated 4 months ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆222Updated this week
- KalamDB — a lightweight, real-time, storage-efficient SQL database. Designed for per-user data isolation and scalable performance — ideal…☆47Updated this week
- Interactive live visualizer for gepa runs☆419May 26, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆51Updated this week
- Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config…☆27Jul 12, 2026Updated 3 weeks ago
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆6,865Jul 9, 2026Updated 3 weeks ago
- top, but for the HTTP endpoints on your host — a live dashboard of the most active plaintext HTTP endpoints, right in the terminal.☆222Jul 10, 2026Updated 3 weeks ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,562May 10, 2026Updated 2 months ago
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated 3 weeks ago
- Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere☆1,356Jul 1, 2026Updated last month
- Adaptive Test-time Learning and Autonomous Specialization☆2,069Updated this week
- Browser System With Zero HTML/CSS/JS☆30Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A minimal hardware-software architecture giving large language models a closed-loop physical embodiment with self-perception loops.☆242Updated this week
- OpenAI's privacy filter NER model architecture implemented in a minimal C++/GGML runtime☆284Jul 2, 2026Updated last month
- Usage indicator extension for pi with footer status bars and /usage command☆15Apr 17, 2026Updated 3 months ago
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆452Jun 22, 2026Updated last month
- ☆35Jul 12, 2026Updated 3 weeks ago
- Development memory that flows between sessions. An MCP server that gives Claude persistent memory about your projects.☆18May 2, 2026Updated 3 months ago
- A full GUI experience on top of llama.cpp: all-knobs model tuning, one-click build/update from upstream, HuggingFace discovery with VRAM-…☆54Updated this week