How much experts do we need to serve a model?
☆152Mar 18, 2026Updated 6 months ago
Alternatives and similar repositories for reap-expert-swap
Users that are interested in reap-expert-swap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bash harness to run Karpathy's autoresearch with Codex CLI in a loop. Includes A/B testing framework for comparing models.☆57Mar 13, 2026Updated 6 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 6 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,822Updated this week
- ☆183Mar 30, 2026Updated 6 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,345Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- FMS Model Optimizer is a framework for developing reduced precision neural network models.☆21Updated this week
- Hermes-native local-first credential broker, scanner, and encrypted vault.☆278Sep 14, 2026Updated 3 weeks ago
- ☆17Oct 31, 2025Updated 11 months ago
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆226Sep 1, 2026Updated last month
- Run LLMs with MLX☆14May 1, 2026Updated 5 months ago
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆24Aug 17, 2026Updated last month
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆525Apr 17, 2026Updated 5 months ago
- Your AI friend right in your browser☆550Sep 17, 2026Updated 3 weeks ago
- FRI Extended for Data Availability: a FRI-based Data Availability Sampling library, written in Rust.☆16May 12, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Headless Matrix WebRTC voice AND video agent — auto-answers calls, bridges audio to any AI agent via PipeWire, optional camera-frame visi…☆20Updated this week
- A cog implementation of Nvidia's Triton server☆18Oct 23, 2024Updated last year
- Pi coding agent extension: llama.cpp provider with dynamic model + context window discovery☆110Sep 24, 2026Updated 2 weeks ago
- Can LLMs be provable computers?☆63Apr 15, 2026Updated 5 months ago
- MCP server that tracks file descriptions across codebases, enabling AI agents to efficiently navigate and understand code through searcha…☆21Feb 4, 2026Updated 8 months ago
- Agent Skill for exploring Obsidian vaults with Enzyme — self-contained, cross-agent compatible☆64Sep 30, 2026Updated last week
- The training codes of Jasper-Token-Compression-600M☆22Nov 19, 2025Updated 10 months ago
- Hermes Dashboard plugins — Honcha Memory, and more☆27Apr 28, 2026Updated 5 months ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆232Oct 3, 2026Updated last week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Updated this week
- ☆38Mar 30, 2026Updated 6 months ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆106Mar 19, 2026Updated 6 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆19Apr 11, 2026Updated 5 months ago
- ☆25Aug 29, 2025Updated last year
- ☆21Aug 31, 2026Updated last month
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- AgentParse is a high-performance parsing library designed to map various structured data formats (such as Pydantic models, JSON, YAML, an…☆18Oct 13, 2025Updated 11 months ago
- A ComfyUI custom node that brings a DAW-style interactive video timeline directly into the node graph. Upload any video, scrub through it…☆19Apr 21, 2026Updated 5 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI serv…☆7,184Updated this week
- ☆54Jan 15, 2026Updated 8 months ago
- experiment loop for ai agents and swarms☆173Mar 29, 2026Updated 6 months ago
- setup the env for vllm users☆16Oct 31, 2023Updated 2 years ago
- Fast Opportunistic Mixture-Of-Experts. From-scratch C/HIP MoE inference with multi-tier caching and cache-aware routing. First ever examp…☆31Mar 31, 2026Updated 6 months ago
- Mikan 🍊: The ZK Friendly DA Layer for Bitcoin L2s☆26Nov 24, 2025Updated 10 months ago
- HermesWorld dashboard plugin for Hermes Agent☆32May 6, 2026Updated 5 months ago