How much experts do we need to serve a model?
☆151Mar 18, 2026Updated 6 months ago
Alternatives and similar repositories for reap-expert-swap
Users that are interested in reap-expert-swap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bash harness to run Karpathy's autoresearch with Codex CLI in a loop. Includes A/B testing framework for comparing models.☆57Mar 13, 2026Updated 6 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 6 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,768Updated this week
- ☆182Mar 30, 2026Updated 5 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,281Sep 3, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The agent that grows with you☆61Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 5 months ago
- FMS Model Optimizer is a framework for developing reduced precision neural network models.☆21Sep 9, 2026Updated last week
- Hermes-native local-first credential broker, scanner, and encrypted vault.☆275Updated this week
- Your AI friend right in your browser☆547Updated this week
- Lightweight autoresearch with guarded self-improvement, simple memory, and Obsidian observability.☆20Jul 27, 2026Updated last month
- MCP server that tracks file descriptions across codebases, enabling AI agents to efficiently navigate and understand code through searcha…☆20Feb 4, 2026Updated 7 months ago
- Agent Skill for exploring Obsidian vaults with Enzyme — self-contained, cross-agent compatible☆62Aug 28, 2026Updated 3 weeks ago
- Hermes Dashboard plugins — Honcha Memory, and more☆27Apr 28, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆24Oct 2, 2025Updated 11 months ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆229Updated this week
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 2 months ago
- Zero-setup bash CLI that downloads full-resolution images from iCloud/Dropbox/Google Photos share links, bridging iPhone screenshots to r…☆43Aug 24, 2026Updated 3 weeks ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆99Mar 19, 2026Updated 6 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆19Apr 11, 2026Updated 5 months ago
- ☆21Aug 31, 2026Updated 2 weeks ago
- Public repo for kernel writing skills in CuTeDSL, Triton, Tilelang and CUDA☆25Jul 7, 2026Updated 2 months ago
- A better way to handle errors. Both data and errors are declared with const, available at the top level, and non-nullable (once the other…☆11Sep 5, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A ComfyUI custom node that brings a DAW-style interactive video timeline directly into the node graph. Upload any video, scrub through it…☆19Apr 21, 2026Updated 4 months ago
- Private IPFS Gateway for better authentication, security, and web3 WebAssembly hosting.☆22Jul 25, 2026Updated last month
- AgentParse is a high-performance parsing library designed to map various structured data formats (such as Pydantic models, JSON, YAML, an…☆18Oct 13, 2025Updated 11 months ago
- experiment loop for ai agents and swarms☆174Mar 29, 2026Updated 5 months ago
- setup the env for vllm users☆16Oct 31, 2023Updated 2 years ago
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.uggested Donat…☆35Jul 20, 2026Updated 2 months ago
- HermesWorld dashboard plugin for Hermes Agent☆32May 6, 2026Updated 4 months ago
- Native Mac OS GUI for Using mlx-lm-lora.☆66Dec 19, 2025Updated 9 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,868Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- HomebrewNLP in JAX flavour for maintable TPU-Training☆50Jan 20, 2024Updated 2 years ago
- Unified API to facilitate usage of pre-trained "perceptor" models, a la CLIP☆39Nov 26, 2022Updated 3 years ago
- Attempt at cog wrapper for nightmareai/real-esrgan for larger images☆16Sep 28, 2023Updated 2 years ago
- ☆17Dec 11, 2025Updated 9 months ago
- ☆17Feb 15, 2026Updated 7 months ago
- Convert Hermes / OpenClaw agent memory (JSONL sessions, MEMORY.md) to markdown for gBrain ingest☆66Apr 13, 2026Updated 5 months ago
- Adaptive Memory is a sophisticated plugin that gives Large Language Models persistent, personalized memory across conversations. It autom…☆21Aug 8, 2026Updated last month