Structured Chain-of-Thought
☆219May 16, 2026Updated 2 months ago
Alternatives and similar repositories for structured-cot
Users that are interested in structured-cot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM☆34Apr 13, 2026Updated 3 months ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,728Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆222Updated this week
- ☆67Jun 4, 2026Updated 2 months ago
- Multi-tab GUI terminal wrapper for AI coding assistants — image paste, Telegram bridge, session persistence.☆18Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 3 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆397Updated this week
- ☆392Apr 16, 2026Updated 3 months ago
- ☆36Apr 25, 2026Updated 3 months ago
- Agent-first knowledge base — wiki over RAG. Plain markdown with wikilinks, background quality agents, and an MCP server for agent navigat…☆27Apr 15, 2026Updated 3 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,579May 10, 2026Updated 3 months ago
- ☆15Apr 26, 2025Updated last year
- Experimental llama.cpp fork for inference research and development☆731Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Give your agent a knowledge graph that compounds.☆34Jun 15, 2026Updated last month
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,638Updated this week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆439Jul 3, 2026Updated last month
- ☆68Jul 29, 2026Updated last week
- Production focused Self-harnessed LM runtime (RLM) that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, output…☆424Updated this week
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆96Mar 19, 2026Updated 4 months ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆79Jul 18, 2026Updated 3 weeks ago
- ☆46Feb 20, 2026Updated 5 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,921Updated this week
- Sequential Monte Carlo Speculative Decoding☆52Jul 25, 2026Updated 2 weeks ago
- Metadspy: The framework for specifying—not programming—language models☆87Jun 18, 2025Updated last year
- Storing the LongCoT-mini results for RLM(GPT-5.2)☆20Apr 26, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,024Updated this week
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 8 months ago
- A plugin for your agentic framework that optimizes code using the GEPA algorithm (Genetic-Pareto LLM-driven search).☆98Apr 28, 2026Updated 3 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆431Jul 25, 2026Updated 2 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 3 months ago
- Evaluation repository of wikipedia index with Dria☆10Mar 14, 2024Updated 2 years ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆305Apr 24, 2026Updated 3 months ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,044Apr 23, 2026Updated 3 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆759Jun 11, 2026Updated 2 months ago
- ☆81Feb 18, 2026Updated 5 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆120Jul 15, 2026Updated 3 weeks ago