Structured Chain-of-Thought
☆220May 16, 2026Updated 4 months ago
Alternatives and similar repositories for structured-cot
Users that are interested in structured-cot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM☆36Apr 13, 2026Updated 5 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,907Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆232Oct 3, 2026Updated last week
- ☆68Jun 4, 2026Updated 4 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 5 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆32May 12, 2026Updated 4 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆429Aug 10, 2026Updated 2 months ago
- direct preference optimization with only 1 model copy :)☆14Oct 2, 2023Updated 3 years ago
- ☆397Apr 16, 2026Updated 5 months ago
- ☆37Apr 25, 2026Updated 5 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,150Aug 18, 2026Updated last month
- ☆15Apr 26, 2025Updated last year
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 5 months ago
- Experimental llama.cpp fork for inference research and development☆887Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Give your agent a knowledge graph that compounds.☆40Jun 15, 2026Updated 3 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,822Updated this week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆478Updated this week
- ☆72Sep 27, 2026Updated 2 weeks ago
- Production focused Self-harnessed LM runtime (RLM) that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, output…☆435Oct 2, 2026Updated last week
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆106Mar 19, 2026Updated 6 months ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆82Jul 18, 2026Updated 2 months ago
- ☆46Feb 20, 2026Updated 7 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, sglang, llama.cpp and other custom engines) and m…☆2,367Updated this week
- Sequential Monte Carlo Speculative Decoding☆53Aug 25, 2026Updated last month
- Metadspy: The framework for specifying—not programming—language models☆87Jun 18, 2025Updated last year
- Storing the LongCoT-mini results for RLM(GPT-5.2)☆20Apr 26, 2026Updated 5 months ago
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- llama.cpp fork with additional SOTA quants and improved performance☆3,290Updated this week
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 10 months ago
- A plugin for your agentic framework that optimizes code using the GEPA algorithm (Genetic-Pareto LLM-driven search).☆99Apr 28, 2026Updated 5 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆475Aug 17, 2026Updated last month
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆20Apr 18, 2026Updated 5 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,054Apr 23, 2026Updated 5 months ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆308Apr 24, 2026Updated 5 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆787Aug 20, 2026Updated last month
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆134Sep 4, 2026Updated last month
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆24Apr 4, 2026Updated 6 months ago
- ☆83Feb 18, 2026Updated 7 months ago