☆117Aug 6, 2026Updated 3 weeks ago
Alternatives and similar repositories for nccl-mesh-plugin
Users that are interested in nccl-mesh-plugin are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 3 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆103Updated this week
- ☆16Jan 15, 2026Updated 7 months ago
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 7 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,202Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆19Oct 25, 2025Updated 10 months ago
- A PyTorch implementation of gradient-free optimization for directly optimizing NDCG (Normalized Discounted Cumulative Gain) in neural inf…☆19Dec 21, 2025Updated 8 months ago
- ☆37Mar 2, 2026Updated 5 months ago
- Edge AI & Offline Vector Search☆15Jan 3, 2026Updated 7 months ago
- ☆35Feb 6, 2026Updated 6 months ago
- Loader extension for tabbyAPI in SillyTavern☆27Jun 30, 2025Updated last year
- ☆14Dec 6, 2023Updated 2 years ago
- The Python Implementation of CRISP: Clustering Multi-Vector Representations for Denoising and Pruning☆27Jul 27, 2025Updated last year
- Hitoku Draft. A context aware macOS local AI assistant. Fully local.☆16Jun 10, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Evolution process to find the best quant tensor weights to build the most optimal GGUF options for an AI model.☆43Aug 22, 2026Updated last week
- A pure and fast NumPy implementation of Mamba with cache support.☆18Jun 16, 2024Updated 2 years ago
- Some benchmark results of small models and quants that fit on DGX Spark☆50Aug 23, 2026Updated last week
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆26Sep 1, 2025Updated 11 months ago
- Docker compose serving stack for DeepSeek v4 Flash DSpark for NVIDIA Spark GB10 system using Aidendle94 image☆34Aug 17, 2026Updated 2 weeks ago
- Mistral Vibe rewritten in Rust by Devstral 2☆21Dec 23, 2025Updated 8 months ago
- a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for eas…☆20Mar 14, 2025Updated last year
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 4 months ago
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆29May 31, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Give AI agents ambitious work without losing the plot. nac is an open-source harness for long-running tasks, using a central orchestrator…☆179Updated this week
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆18Aug 6, 2026Updated 3 weeks ago
- llama-benchy - llama-bench style benchmarking tool for all backends☆682Jul 10, 2026Updated last month
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- AI-driven automated threat analysis pipeline that routes files, URLs, IPs, domains, or images through specialized security analyzers and …☆16Mar 9, 2026Updated 5 months ago
- A simple speech-to-text and text-to-speech AI chatbot that can be run fully offline.☆45Jan 28, 2024Updated 2 years ago
- ONNX speech pipeline library for ASR, diarization, VAD, and denoising☆20Updated this week
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆105Jul 28, 2026Updated last month
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆483Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆52Nov 17, 2025Updated 9 months ago
- Proves that larger context windows don't fix RAG on structured data — they make wrong answers harder to detect. Then solves it with a que…☆15Updated this week
- Simple, Fast, Parallel Huggingface GGML model downloader written in python☆24Jul 26, 2023Updated 3 years ago
- Code for the ICML 2025 Paper "Product of Experts with LLMs: Boosting Performance on ARC is a Matter of Perspective"☆56Nov 9, 2025Updated 9 months ago
- ☆18Apr 21, 2026Updated 4 months ago
- A systematic empirical study of self-verification strategies in agentic coding harnesses☆29Mar 4, 2026Updated 5 months ago
- The most feature-complete local AI workstation. Multi-GPU inference, integrated Stable Diffusion + ADetailer, voice cloning, research-gra…☆67Jul 28, 2026Updated last month