Speculative Decoding Implementations: MTP, EAGLE-3, Medusa-1, PARD, Draft Models, N-gram and Suffix Decoding from scratch
☆15May 2, 2026Updated 5 months ago
Alternatives and similar repositories for Speculative-Decoding
Users that are interested in Speculative-Decoding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A collection of optimized ComfyUI-based cloud inference endpoints, built on ComfyDeploy and Modal☆16Nov 5, 2024Updated last year
- Multi-agent orchestration framework for AI applications - build, deploy, and manage AI agents across the full lifecycle with Forge, Conve…☆33Mar 28, 2026Updated 6 months ago
- ☆19Mar 19, 2026Updated 6 months ago
- ☆16Feb 3, 2026Updated 8 months ago
- *NIX SHELL with Local AI/LLM integration☆26Feb 26, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 9 months ago
- A text-grid web renderer for AI agents — see the web without screenshots☆104Mar 10, 2026Updated 6 months ago
- Your AI forgets everything between sessions. SAME fixes that. Local-first, no API keys, single binary.☆23Sep 3, 2026Updated last month
- Tool schemas get re-prefilled on every LLM request. ContextCache compiles them into a KV cache once, stores it on disk and reuses it acro…☆22Sep 16, 2026Updated 2 weeks ago
- Technical Analysis Library using Pandas (Modin for speedup) (Python)☆11Jun 24, 2019Updated 7 years ago
- A bytebot variant that uses Holo 1.5 7b to control the desktop☆25Nov 4, 2025Updated 10 months ago
- Local LLM Search Index for RAG☆20May 5, 2026Updated 4 months ago
- A yabai status bar widget for Übersicht☆18Dec 20, 2020Updated 5 years ago
- ☆18Dec 1, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An AI assistant for PCs powered by Meta's LLaMA3 using Hugging Face, runs on voice recognition, text-to-speech. Send messages, voice/vide…☆20Jun 6, 2024Updated 2 years ago
- ☆72Jul 3, 2026Updated 3 months ago
- ☆18Apr 1, 2024Updated 2 years ago
- Lightweight API Specification for Intelligent Systems☆16Sep 8, 2026Updated 3 weeks ago
- CLI Orchestrator for AI Agents☆25Feb 18, 2026Updated 7 months ago
- ☆15Mar 18, 2026Updated 6 months ago
- Git-native shared memory for AI CLI agents☆38Jun 4, 2026Updated 3 months ago
- ☆12Sep 2, 2021Updated 5 years ago
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- a terminal based GPS tool.☆17Dec 8, 2025Updated 9 months ago
- A simple model context protocol (MCP) server that allows Claude Desktop or other MCP aware clients to run Bash commands on your local mac…☆35Apr 14, 2025Updated last year
- Parser and printer for the opam file syntax☆16Jul 3, 2025Updated last year
- SIMBL plugin for macOS that hides traffic lights buttons☆15Jun 28, 2024Updated 2 years ago
- ☆10May 9, 2019Updated 7 years ago
- Write data to files split by topic and rolled over on size or a timeout, files can be compressed using lzo, snappy or gzip☆11Jul 12, 2021Updated 5 years ago
- A multi-agent AI chat running on Convex☆18Apr 4, 2025Updated last year
- A modern web-based management interface for Proxmox VE with multi-cluster support and focus on desktop and mobile experience.☆62Updated this week
- DocFinder is a local-first indexing and searching documents using semantic embeddings stored in SQLite. Everything runs on your machine, …☆33Sep 26, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆25Feb 16, 2026Updated 7 months ago
- ☆20Jan 3, 2026Updated 9 months ago
- Packaged externs for the Node.js standard library + utilities☆12Oct 24, 2017Updated 8 years ago
- A pod around buddy core (Cryptographic Api for Clojure).☆17Jun 6, 2024Updated 2 years ago
- ☆19Jun 11, 2025Updated last year
- TeGere! = Behave! — a Gherkin library for Clojure☆13Oct 12, 2023Updated 2 years ago
- Nano vLLM v1 engine☆15Aug 6, 2025Updated last year