Fast, lossless LLM inference via dual-view diffusion decoding.
☆460May 18, 2026Updated 2 months ago
Alternatives and similar repositories for orthrus
Users that are interested in orthrus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆440Jun 22, 2026Updated 3 weeks ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆794Updated this week
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆121Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,504May 10, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows☆2,189Updated this week
- 26m agentic model for tiny devices☆3,236Updated this week
- A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD),…☆21Updated this week
- Coding Agent singularly focused efficiency and context curation. Reduces API costs by 50-80% vs other agent AND improves the code quality…☆1,413Updated this week
- An open-source, self-hosted mini Bloomberg Terminal for AI agents, exposed as an MCP server — SEC filings, institutional holdings, inside…☆183Updated this week
- A local-first web search agent☆29Jun 20, 2026Updated last month
- lowfat - slim your command output. strips noise, saves tokens.☆560Jul 8, 2026Updated 2 weeks ago
- ☆180Apr 27, 2026Updated 2 months ago
- An embeddable cross-platform AI agent library written in C. Cloud and local LLMs, tool calling, long-term memory, voice, sessions, resear…☆100May 21, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆163Jun 27, 2026Updated 3 weeks ago
- The Rex Programming Language☆18Updated this week
- ModelCypher - Decipher the high dimensional geometry of LLMs. An open source x-ray into LLM representation structure.☆25Jul 14, 2026Updated last week
- Adaptive Test-time Learning and Autonomous Specialization☆2,064Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆2,943Updated this week
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Run models too big for your Mac's memory☆662Apr 8, 2026Updated 3 months ago
- Automated parameter sweep pipeline for finding optimal sampling settings for any local LLM on quantized weights☆18Feb 21, 2026Updated 5 months ago
- Fast and Accurate Code Search for Agents. Uses ~98% fewer tokens than grep+read☆5,665Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Self-hosted email gateway between your apps and a transactional mail provider (Postmark, Resend, Mailgun, AWS SES, or outbound-SMTP). Thr…☆205Jul 11, 2026Updated last week
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,761Updated this week
- Ultra-minimal personal AI agent: starts small, self-modifies its code live, adapts by writing exactly the code & features you need☆218Feb 27, 2026Updated 4 months ago
- I replicated Ng's RYS method and found that duplicating 3 specific layers in Qwen2.5-32B boosts reasoning by 17% and duplicating layers 1…☆242Mar 20, 2026Updated 4 months ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆22Mar 6, 2026Updated 4 months ago
- [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models☆1,021Jul 10, 2025Updated last year
- The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.☆16Updated this week
- An open reference standard for jiggle physics: weight-painted regions + damped spring bones, one rule (vertex += weight * boneJiggle). Po…☆35May 31, 2026Updated last month
- HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.☆1,712Jun 17, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 7 months ago
- Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~2x upstream prefill, ~1.5x decode, DSpar…☆64Updated this week
- A novel hybrid AI architecture leveraging Titan's-like memory and HRM-like reasoning☆26Updated this week
- ☆75Jun 10, 2026Updated last month
- State machine guardrails for AI agents☆417Updated this week
- Code for the manuscript "A self-supervised multi-layer network of Rectified Spectral Units (ReSUs)" submitted to NeurIPS☆25Apr 6, 2026Updated 3 months ago
- Run a 120B-parameter MoE (60 GB) on a 12 GB phone. CPU-only, lossless, on stock llama.cpp☆32Updated this week