Structured Chain-of-Thought
☆219May 16, 2026Updated 2 months ago
Alternatives and similar repositories for structured-cot
Users that are interested in structured-cot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM☆34Apr 13, 2026Updated 3 months ago
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆221Jul 6, 2026Updated 2 weeks ago
- ☆68Jun 4, 2026Updated last month
- Multi-tab GUI terminal wrapper for AI coding assistants — image paste, Telegram bridge, session persistence.☆18Jul 15, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 3 months ago
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆27May 12, 2026Updated 2 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆376Jun 16, 2026Updated last month
- direct preference optimization with only 1 model copy :)☆14Oct 2, 2023Updated 2 years ago
- ☆389Apr 16, 2026Updated 3 months ago
- ☆36Apr 25, 2026Updated 2 months ago
- Agent-first knowledge base — wiki over RAG. Plain markdown with wikilinks, background quality agents, and an MCP server for agent navigat…☆27Apr 15, 2026Updated 3 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,504May 10, 2026Updated 2 months ago
- LLAMA Turboquant implementation with CUDA support☆704Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Give your agent a knowledge graph that compounds.☆32Jun 15, 2026Updated last month
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,481Updated this week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆427Jul 3, 2026Updated 2 weeks ago
- ☆68Updated this week
- Production focused Self-harnessed LM runtime (RLM) that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, output…☆412Updated this week
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆96Mar 19, 2026Updated 4 months ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆72Updated this week
- ☆46Feb 20, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆34May 26, 2026Updated last month
- REAM: Merging Improves Pruning of Experts in LLMs☆22Apr 16, 2026Updated 3 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,761Updated this week
- Sequential Monte Carlo Speculative Decoding☆52Updated this week
- Metadspy: The framework for specifying—not programming—language models☆87Jun 18, 2025Updated last year
- Storing the LongCoT-mini results for RLM(GPT-5.2)☆20Apr 26, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,943Updated this week
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- Find the hidden meaning of LLMs☆41Nov 13, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A plugin for your agentic framework that optimizes code using the GEPA algorithm (Genetic-Pareto LLM-driven search).☆96Apr 28, 2026Updated 2 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆397May 29, 2026Updated last month
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 3 months ago
- Evaluation repository of wikipedia index with Dria☆10Mar 14, 2024Updated 2 years ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,037Apr 23, 2026Updated 2 months ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆304Apr 24, 2026Updated 2 months ago
- ☆81Feb 18, 2026Updated 5 months ago