Fast, lossless LLM inference via dual-view diffusion decoding.
☆477Aug 12, 2026Updated 2 weeks ago
Alternatives and similar repositories for orthrus
Users that are interested in orthrus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆480Jun 22, 2026Updated 2 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆137Jul 25, 2026Updated last month
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,011Aug 18, 2026Updated last week
- A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows☆2,229Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆9,776Updated this week
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆21Aug 24, 2026Updated last week
- Coding Agent singularly focused efficiency and context curation. Reduces API costs by 50-80% vs other agent AND improves the code quality…☆1,470Updated this week
- Self-hosted, open-source financial data MCP server for AI agents — SEC filings, XBRL financials, 13F holdings, insider & congressional tr…☆203Updated this week
- A local-first web search agent☆30Jun 20, 2026Updated 2 months ago
- lowfat - slim your command output. strips noise, saves tokens.☆570Aug 17, 2026Updated 2 weeks ago
- ☆184Apr 27, 2026Updated 4 months ago
- ☆20Apr 8, 2025Updated last year
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆178Aug 9, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- An embeddable cross-platform AI agent library written in C. Cloud and local LLMs, tool calling, long-term memory, voice, sessions, resear…☆119May 21, 2026Updated 3 months ago
- The Rex Programming Language☆41Aug 24, 2026Updated last week
- Adaptive Test-time Learning and Autonomous Specialization☆2,079Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Run models too big for your Mac's memory☆668Updated this week
- Automated parameter sweep pipeline for finding optimal sampling settings for any local LLM on quantized weights☆22Feb 21, 2026Updated 6 months ago
- Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read☆5,968Updated this week
- Self-hosted email gateway between your apps and a transactional mail provider (Postmark, Resend, Mailgun, AWS SES, or outbound-SMTP). Thr…☆210Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Ultra-minimal personal AI agent: starts small, self-modifies its code live, adapts by writing exactly the code & features you need☆217Feb 27, 2026Updated 6 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,122Updated this week
- I replicated Ng's RYS method and found that duplicating 3 specific layers in Qwen2.5-32B boosts reasoning by 17% and duplicating layers 1…☆242Mar 20, 2026Updated 5 months ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆23Mar 6, 2026Updated 5 months ago
- [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models☆1,030Jul 10, 2025Updated last year
- An open reference standard for jiggle physics: weight-painted regions + damped spring bones, one rule (vertex += weight * boneJiggle). Po…☆39May 31, 2026Updated 3 months ago
- HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.☆1,920Jun 17, 2026Updated 2 months ago
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpar…☆384Updated this week
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A novel hybrid AI architecture leveraging Titan's-like memory and HRM-like reasoning☆26Updated this week
- ☆78Updated this week
- State machine guardrails for AI agents☆489Aug 22, 2026Updated last week
- General-purpose psychology agent (Caude Code): collegial mentor, specialized sub-agents, and a consensus-or-parsimony adversarial evaluat…☆20May 1, 2026Updated 3 months ago
- The best Claude Code that $200 can buy☆269Apr 6, 2026Updated 4 months ago
- A harness optimized to smaller LLMs☆2,515Updated this week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 4 months ago