Fast, lossless LLM inference via dual-view diffusion decoding.
☆481Aug 12, 2026Updated last month
Alternatives and similar repositories for orthrus
Users that are interested in orthrus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆497Jun 22, 2026Updated 2 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,101Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,868Updated this week
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated last month
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,105Aug 18, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows☆2,248Sep 1, 2026Updated 2 weeks ago
- Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smar…☆11,854Updated this week
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆20Updated this week
- Coding Agent singularly focused efficiency and context curation. Reduces API costs by 50-80% vs other agent AND improves the code quality…☆1,506Updated this week
- A local-first web search agent☆30Jun 20, 2026Updated 3 months ago
- lowfat - slim your command output. strips noise, saves tokens.☆574Updated this week
- ☆184Apr 27, 2026Updated 4 months ago
- ☆20Apr 8, 2025Updated last year
- An embeddable cross-platform AI agent library written in C. Cloud and local LLMs, tool calling, long-term memory, voice, sessions, resear…☆123May 21, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The Rex Programming Language☆41Aug 24, 2026Updated 3 weeks ago
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆225Aug 9, 2026Updated last month
- Adaptive Test-time Learning and Autonomous Specialization☆2,091Updated this week
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,249Updated this week
- Run models too big for your Mac's memory☆668Aug 27, 2026Updated 3 weeks ago
- Automated parameter sweep pipeline for finding optimal sampling settings for any local LLM on quantized weights☆22Feb 21, 2026Updated 6 months ago
- Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read☆6,113Updated this week
- Ultra-minimal personal AI agent: starts small, self-modifies its code live, adapts by writing exactly the code & features you need☆220Feb 27, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- I replicated Ng's RYS method and found that duplicating 3 specific layers in Qwen2.5-32B boosts reasoning by 17% and duplicating layers 1…☆242Mar 20, 2026Updated 6 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,277Updated this week
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆23Mar 6, 2026Updated 6 months ago
- [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models☆1,033Jul 10, 2025Updated last year
- An open reference standard for jiggle physics: weight-painted regions + damped spring bones, one rule (vertex += weight * boneJiggle). Po…☆40Sep 7, 2026Updated last week
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 9 months ago
- HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.☆2,052Sep 4, 2026Updated 2 weeks ago
- Cross-platform Rust daemon for scheduling AI-agent routines via UI, REST, and MCP.☆53Updated this week
- A novel hybrid AI architecture leveraging Titan's-like memory and HRM-like reasoning☆27Aug 28, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpar…☆402Aug 27, 2026Updated 3 weeks ago
- ☆78Updated this week
- State machine guardrails for AI agents☆493Updated this week
- Code for the manuscript "A self-supervised multi-layer network of Rectified Spectral Units (ReSUs)" submitted to NeurIPS☆25Apr 6, 2026Updated 5 months ago
- The best Claude Code that $200 can buy☆268Apr 6, 2026Updated 5 months ago
- A harness optimized to smaller LLMs☆2,606Updated this week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆322Apr 19, 2026Updated 5 months ago