☆390Apr 16, 2026Updated 3 months ago
Alternatives and similar repositories for ddtree
Users that are interested in ddtree are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,543May 10, 2026Updated 2 months ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆757Jun 11, 2026Updated last month
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 3 months ago
- Exact speculative decoding on Apple Silicon, powered by MLX.☆379Apr 20, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆171Jun 27, 2026Updated last month
- Fast LLM speculative inference server for consumer hardware.☆2,689Updated this week
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,018Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆660Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆326Jul 1, 2026Updated 3 weeks ago
- A lightweight inference engine supporting speculative speculative decoding (SSD).☆975May 10, 2026Updated 2 months ago
- DMax: Aggressive Parallel Decoding for dLLMs☆128Jul 5, 2026Updated 3 weeks ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆830Jul 14, 2026Updated 2 weeks ago
- Sequential Monte Carlo Speculative Decoding☆52Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 2 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,041Apr 23, 2026Updated 3 months ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,481Feb 20, 2026Updated 5 months ago
- A PyTorch native library for training speculative decoding models☆208Updated this week
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.☆211Mar 18, 2026Updated 4 months ago
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆124Updated this week
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆53Jun 28, 2026Updated last month
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆117Jul 15, 2026Updated 2 weeks ago
- DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms☆6,803Jul 9, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- TokenSpeed is a speed-of-light LLM inference engine.☆1,717Updated this week
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆463Jul 14, 2026Updated 2 weeks ago
- ☆7,003Jul 20, 2026Updated last week
- PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM☆34Apr 13, 2026Updated 3 months ago
- Structured Chain-of-Thought☆219May 16, 2026Updated 2 months ago
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆174Feb 27, 2026Updated 5 months ago
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆275Jun 3, 2026Updated last month
- Compact, local-first neuro-symbolic assistant in Rust with a 200 MiB Bitwork cognitive model, exact reasoning tools, governed memory, and…☆42Updated this week
- Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU pl…☆24Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"☆1,064May 30, 2026Updated last month
- Official Implementation of DART (DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference).☆64Feb 8, 2026Updated 5 months ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆402Apr 22, 2025Updated last year
- FlashKDA: high-performance Kimi Delta Attention kernels☆835Updated this week
- ☆72Jun 3, 2026Updated last month
- high-performance linear attention kernel library built on TileLang☆616Updated this week
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆22May 25, 2026Updated 2 months ago