How much experts do we need to serve a model?
☆150Mar 18, 2026Updated 5 months ago
Alternatives and similar repositories for reap-expert-swap
Users that are interested in reap-expert-swap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,745Updated this week
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,264Aug 19, 2026Updated last week
- The agent that grows with you☆62Aug 7, 2026Updated 3 weeks ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FMS Model Optimizer is a framework for developing reduced precision neural network models.☆21Jun 24, 2026Updated 2 months ago
- Hermes-native local-first credential broker, scanner, and encrypted vault.☆261Updated this week
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆21Aug 17, 2026Updated 2 weeks ago
- ☆17Oct 31, 2025Updated 9 months ago
- Run LLMs with MLX☆15May 1, 2026Updated 3 months ago
- Bounded, inspectable LLM inference pipelines from declared YAML — runs offline against Ollama or any OpenAI-compatible local server, emit…☆46Jul 7, 2026Updated last month
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆490Apr 17, 2026Updated 4 months ago
- Your AI friend right in your browser☆542Apr 24, 2026Updated 4 months ago
- FRI Extended for Data Availability: a FRI-based Data Availability Sampling library, written in Rust.☆16May 12, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pi coding agent extension: llama.cpp provider with dynamic model + context window discovery☆97Jul 6, 2026Updated last month
- Can LLMs be provable computers?☆63Apr 15, 2026Updated 4 months ago
- MCP server that tracks file descriptions across codebases, enabling AI agents to efficiently navigate and understand code through searcha…☆20Feb 4, 2026Updated 6 months ago
- Agent Skill for exploring Obsidian vaults with Enzyme — self-contained, cross-agent compatible☆63Updated this week
- Hermes Dashboard plugins — Honcha Memory, and more☆28Apr 28, 2026Updated 4 months ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆230Updated this week
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 4 months ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 2 months ago
- Zero-setup bash CLI that downloads full-resolution images from iCloud/Dropbox/Google Photos share links, bridging iPhone screenshots to r…☆43Aug 24, 2026Updated last week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆98Mar 19, 2026Updated 5 months ago
- ☆21Jan 15, 2026Updated 7 months ago
- Issue tracker for https://alphaxiv.org☆26Oct 13, 2025Updated 10 months ago
- ASHIGARU is a terminal-based interface & system powered by React & Ink that combines the power of modern web technologies with the effici…☆48Jan 2, 2026Updated 7 months ago
- Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.☆5,145Updated this week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- A ComfyUI custom node that brings a DAW-style interactive video timeline directly into the node graph. Upload any video, scrub through it…☆18Apr 21, 2026Updated 4 months ago
- Hermes Relay for Chrome gives Hermes Agent a direct browser surface for page context, capture, watchlists, and AI handoff.☆22May 4, 2026Updated 3 months ago
- Private IPFS Gateway for better authentication, security, and web3 WebAssembly hosting.☆22Jul 25, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- AgentParse is a high-performance parsing library designed to map various structured data formats (such as Pydantic models, JSON, YAML, an…☆18Oct 13, 2025Updated 10 months ago
- ☆55Jan 15, 2026Updated 7 months ago
- experiment loop for ai agents and swarms☆174Mar 29, 2026Updated 5 months ago
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.uggested Donat…☆33Jul 20, 2026Updated last month
- Core programs of the Exponent protocol.☆34Aug 18, 2026Updated last week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- Unified API to facilitate usage of pre-trained "perceptor" models, a la CLIP☆39Nov 26, 2022Updated 3 years ago