Structured Chain-of-Thought
☆220May 16, 2026Updated 3 months ago
Alternatives and similar repositories for structured-cot
Users that are interested in structured-cot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM☆35Apr 13, 2026Updated 4 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆230Updated this week
- ☆67Jun 4, 2026Updated 2 months ago
- Multi-tab GUI terminal wrapper for AI coding assistants — image paste, Telegram bridge, session persistence.☆18Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 4 months ago
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆28May 12, 2026Updated 3 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆415Aug 10, 2026Updated 3 weeks ago
- direct preference optimization with only 1 model copy :)☆14Oct 2, 2023Updated 2 years ago
- ☆395Apr 16, 2026Updated 4 months ago
- ☆36Apr 25, 2026Updated 4 months ago
- Agent-first knowledge base — wiki over RAG. Plain markdown with wikilinks, background quality agents, and an MCP server for agent navigat…☆27Apr 15, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,011Aug 18, 2026Updated last week
- ☆15Apr 26, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Experimental llama.cpp fork for inference research and development☆789Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- Give your agent a knowledge graph that compounds.☆36Jun 15, 2026Updated 2 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,745Updated this week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆454Jul 3, 2026Updated last month
- ☆68Updated this week
- Production focused Self-harnessed LM runtime (RLM) that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, output…☆429Aug 5, 2026Updated 3 weeks ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆98Mar 19, 2026Updated 5 months ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆80Jul 18, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆46Feb 20, 2026Updated 6 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 3 months ago
- REAM: Merging Improves Pruning of Experts in LLMs☆25Apr 16, 2026Updated 4 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,122Updated this week
- Metadspy: The framework for specifying—not programming—language models☆87Jun 18, 2025Updated last year
- Storing the LongCoT-mini results for RLM(GPT-5.2)☆20Apr 26, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A plugin for your agentic framework that optimizes code using the GEPA algorithm (Genetic-Pareto LLM-driven search).☆100Apr 28, 2026Updated 4 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆460Aug 17, 2026Updated 2 weeks ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 4 months ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆305Apr 24, 2026Updated 4 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,042Apr 23, 2026Updated 4 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆776Aug 20, 2026Updated last week
- ☆81Feb 18, 2026Updated 6 months ago