How much experts do we need to serve a model?
☆153Mar 18, 2026Updated 4 months ago
Alternatives and similar repositories for reap-expert-swap
Users that are interested in reap-expert-swap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bash harness to run Karpathy's autoresearch with Codex CLI in a loop. Includes A/B testing framework for comparing models.☆52Mar 13, 2026Updated 4 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,476Updated this week
- ☆176Mar 30, 2026Updated 3 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆842Jan 21, 2026Updated 5 months ago
- The agent that grows with you☆64Jul 9, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Hermes-native local-first credential broker, scanner, and encrypted vault.☆237Updated this week
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆17Jul 5, 2026Updated 2 weeks ago
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆52Updated this week
- ☆17Oct 31, 2025Updated 8 months ago
- Run LLMs with MLX☆15May 1, 2026Updated 2 months ago
- Bounded, inspectable LLM inference pipelines from declared YAML — runs offline against Ollama or any OpenAI-compatible local server, emit…☆44Jul 7, 2026Updated last week
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆445Apr 17, 2026Updated 3 months ago
- Your AI friend right in your browser☆541Apr 24, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Pi coding agent extension: llama.cpp provider with dynamic model + context window discovery☆64Jul 6, 2026Updated 2 weeks ago
- ☆17Oct 6, 2025Updated 9 months ago
- Can LLMs be provable computers?☆63Apr 15, 2026Updated 3 months ago
- Lightweight autoresearch with guarded self-improvement, simple memory, and Obsidian observability.☆20Jun 26, 2026Updated 3 weeks ago
- MCP server that tracks file descriptions across codebases, enabling AI agents to efficiently navigate and understand code through searcha…☆18Feb 4, 2026Updated 5 months ago
- Agent Skill for exploring Obsidian vaults with Enzyme — self-contained, cross-agent compatible☆63Jul 6, 2026Updated 2 weeks ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆221Jul 6, 2026Updated 2 weeks ago
- The training codes of Jasper-Token-Compression-600M☆20Nov 19, 2025Updated 8 months ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆42Jun 28, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Zero-setup bash CLI that downloads full-resolution images from iCloud/Dropbox/Google Photos share links, bridging iPhone screenshots to r…☆40Jun 8, 2026Updated last month
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆96Mar 19, 2026Updated 4 months ago
- ☆24Aug 29, 2025Updated 10 months ago
- A ComfyUI custom node that brings a DAW-style interactive video timeline directly into the node graph. Upload any video, scrub through it…☆18Apr 21, 2026Updated 3 months ago
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆17Jul 13, 2026Updated last week
- A better way to handle errors. Both data and errors are declared with const, available at the top level, and non-nullable (once the other…☆11Sep 5, 2024Updated last year
- ☆55Jan 15, 2026Updated 6 months ago
- Hermes Relay for Chrome gives Hermes Agent a direct browser surface for page context, capture, watchlists, and AI handoff.☆22May 4, 2026Updated 2 months ago
- experiment loop for ai agents and swarms☆175Mar 29, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.☆29Updated this week
- setup the env for vllm users☆16Oct 31, 2023Updated 2 years ago
- Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.☆3,318Updated this week
- Native Mac OS GUI for Using mlx-lm-lora.☆64Dec 19, 2025Updated 7 months ago
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- HomebrewNLP in JAX flavour for maintable TPU-Training☆50Jan 20, 2024Updated 2 years ago
- Convert Hermes / OpenClaw agent memory (JSONL sessions, MEMORY.md) to markdown for gBrain ingest☆59Apr 13, 2026Updated 3 months ago