Fast, lossless LLM inference via dual-view diffusion decoding.
☆472May 18, 2026Updated 2 months ago
Alternatives and similar repositories for orthrus
Users that are interested in orthrus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. …☆452Jun 22, 2026Updated last month
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆852Updated this week
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,728Updated this week
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆129Jul 25, 2026Updated 2 weeks ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,579May 10, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows☆2,213Updated this week
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆3,402Updated this week
- A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD),…☆21Aug 1, 2026Updated last week
- Coding Agent singularly focused efficiency and context curation. Reduces API costs by 50-80% vs other agent AND improves the code quality…☆1,451Updated this week
- Self-hosted, open-source financial data MCP server for AI agents — SEC filings, XBRL financials, 13F holdings, insider & congressional tr…☆195Updated this week
- A local-first web search agent☆29Jun 20, 2026Updated last month
- lowfat - slim your command output. strips noise, saves tokens.☆564Jul 8, 2026Updated last month
- ☆180Apr 27, 2026Updated 3 months ago
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting☆176Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- An embeddable cross-platform AI agent library written in C. Cloud and local LLMs, tool calling, long-term memory, voice, sessions, resear…☆114May 21, 2026Updated 2 months ago
- The Rex Programming Language☆17Updated this week
- Adaptive Test-time Learning and Autonomous Specialization☆2,073Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,024Updated this week
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Run models too big for your Mac's memory☆666Apr 8, 2026Updated 4 months ago
- Automated parameter sweep pipeline for finding optimal sampling settings for any local LLM on quantized weights☆19Feb 21, 2026Updated 5 months ago
- Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read☆5,847Updated this week
- Self-hosted email gateway between your apps and a transactional mail provider (Postmark, Resend, Mailgun, AWS SES, or outbound-SMTP). Thr…☆209Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Ultra-minimal personal AI agent: starts small, self-modifies its code live, adapts by writing exactly the code & features you need☆218Feb 27, 2026Updated 5 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,921Updated this week
- I replicated Ng's RYS method and found that duplicating 3 specific layers in Qwen2.5-32B boosts reasoning by 17% and duplicating layers 1…☆242Mar 20, 2026Updated 4 months ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆22Mar 6, 2026Updated 5 months ago
- An open reference standard for jiggle physics: weight-painted regions + damped spring bones, one rule (vertex += weight * boneJiggle). Po…☆36May 31, 2026Updated 2 months ago
- HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.☆1,813Jun 17, 2026Updated last month
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 8 months ago
- A novel hybrid AI architecture leveraging Titan's-like memory and HRM-like reasoning☆26Aug 2, 2026Updated last week
- ☆77Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- State machine guardrails for AI agents☆482Updated this week
- Code for the manuscript "A self-supervised multi-layer network of Rectified Spectral Units (ReSUs)" submitted to NeurIPS☆25Apr 6, 2026Updated 4 months ago
- General-purpose psychology agent (Caude Code): collegial mentor, specialized sub-agents, and a consensus-or-parsimony adversarial evaluat…☆18May 1, 2026Updated 3 months ago
- The best Claude Code that $200 can buy☆269Apr 6, 2026Updated 4 months ago
- A harness optimized to smaller LLMs☆2,343Jul 31, 2026Updated last week
- TurboQuant WASM SIMD vector compression — 3 bits/dim with fast dot product. Requires relaxed SIMD (Chrome 114+, Firefox 128+, Safari 18+,…☆320Apr 19, 2026Updated 3 months ago
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆509Updated this week