How much experts do we need to serve a model?
☆152Mar 18, 2026Updated 4 months ago
Alternatives and similar repositories for reap-expert-swap
Users that are interested in reap-expert-swap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bash harness to run Karpathy's autoresearch with Codex CLI in a loop. Includes A/B testing framework for comparing models.☆53Mar 13, 2026Updated 4 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,621Updated this week
- ☆179Mar 30, 2026Updated 4 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆847Jan 21, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The agent that grows with you☆62Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- FMS Model Optimizer is a framework for developing reduced precision neural network models.☆21Jun 24, 2026Updated last month
- Hermes-native local-first credential broker, scanner, and encrypted vault.☆259Updated this week
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆18Jul 5, 2026Updated last month
- Run LLMs with MLX☆15May 1, 2026Updated 3 months ago
- Bounded, inspectable LLM inference pipelines from declared YAML — runs offline against Ollama or any OpenAI-compatible local server, emit…☆45Jul 7, 2026Updated last month
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆474Apr 17, 2026Updated 3 months ago
- Your AI friend right in your browser☆542Apr 24, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Headless Matrix WebRTC voice AND video agent — auto-answers calls, bridges audio to any AI agent via PipeWire, optional camera-frame visi…☆18Jun 28, 2026Updated last month
- Pi coding agent extension: llama.cpp provider with dynamic model + context window discovery☆82Jul 6, 2026Updated last month
- A cog implementation of Nvidia's Triton server☆18Oct 23, 2024Updated last year
- ☆17Oct 6, 2025Updated 10 months ago
- Can LLMs be provable computers?☆63Apr 15, 2026Updated 3 months ago
- MCP server that tracks file descriptions across codebases, enabling AI agents to efficiently navigate and understand code through searcha…☆18Feb 4, 2026Updated 6 months ago
- Agent Skill for exploring Obsidian vaults with Enzyme — self-contained, cross-agent compatible☆64Updated this week
- The training codes of Jasper-Token-Compression-600M☆20Nov 19, 2025Updated 8 months ago
- ☆24Oct 2, 2025Updated 10 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆222Updated this week
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 3 months ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆42Jun 28, 2026Updated last month
- ☆36Mar 30, 2026Updated 4 months ago
- Zero-setup bash CLI that downloads full-resolution images from iCloud/Dropbox/Google Photos share links, bridging iPhone screenshots to r…☆41Jun 8, 2026Updated 2 months ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆96Mar 19, 2026Updated 4 months ago
- ☆21Jan 15, 2026Updated 6 months ago
- Public repo for kernel writing skills in CuTeDSL, Triton, Tilelang and CUDA☆24Jul 7, 2026Updated last month
- Issue tracker for https://alphaxiv.org☆24Oct 13, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.☆4,071Jul 31, 2026Updated last week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 3 weeks ago
- A better way to handle errors. Both data and errors are declared with const, available at the top level, and non-nullable (once the other…☆11Sep 5, 2024Updated last year
- A ComfyUI custom node that brings a DAW-style interactive video timeline directly into the node graph. Upload any video, scrub through it…☆18Apr 21, 2026Updated 3 months ago
- AgentParse is a high-performance parsing library designed to map various structured data formats (such as Pydantic models, JSON, YAML, an…☆18Oct 13, 2025Updated 9 months ago
- ☆55Jan 15, 2026Updated 6 months ago
- experiment loop for ai agents and swarms☆175Mar 29, 2026Updated 4 months ago