☆110Mar 6, 2026Updated 4 months ago
Alternatives and similar repositories for nccl-mesh-plugin
Users that are interested in nccl-mesh-plugin are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆124Feb 22, 2026Updated 5 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆98Updated this week
- Local-first RAG application for technical documentation and research papers☆28Jun 26, 2026Updated last month
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 6 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆1,944Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆19Oct 25, 2025Updated 9 months ago
- A PyTorch implementation of gradient-free optimization for directly optimizing NDCG (Normalized Discounted Cumulative Gain) in neural inf…☆19Dec 21, 2025Updated 7 months ago
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated 3 weeks ago
- ☆18Jul 1, 2025Updated last year
- ☆35Feb 6, 2026Updated 5 months ago
- ☆11Sep 18, 2023Updated 2 years ago
- Workspace manager for AI coding shells and agents.☆20Updated this week
- "Pacha" TUI (Text User Interface) is a JavaScript application that utilizes the "blessed" library. It serves as a frontend for llama.cpp …☆37Aug 3, 2023Updated 3 years ago
- Loader extension for tabbyAPI in SillyTavern☆27Jun 30, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Mixed-precision numerics benchmarks in Rust and Python - covering GEMMs, SYRKs, DOTs, and higher-level BLAS and LAPACK-style functionalit…☆32Jul 23, 2026Updated last week
- ☆14Dec 6, 2023Updated 2 years ago
- wubuwizard: The best inference engine☆28Updated this week
- The Python Implementation of CRISP: Clustering Multi-Vector Representations for Denoising and Pruning☆27Jul 27, 2025Updated last year
- Production-ready ternary quantized (1.58-bit) Rust code generation model with mHC-lite, MaxRL training, and comprehensive benchmarking☆22Updated this week
- This is the modified version of llama2.c LLM inference app ported to run on 32-bit capable DOS machines.☆30May 23, 2025Updated last year
- Evolution process to find the best quant tensor weights to build the most optimal GGUF options for an AI model.☆39May 13, 2026Updated 2 months ago
- A pure and fast NumPy implementation of Mamba with cache support.☆18Jun 16, 2024Updated 2 years ago
- Some benchmark results of small models and quants that fit on DGX Spark☆49Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Mistral Vibe rewritten in Rust by Devstral 2☆21Dec 23, 2025Updated 7 months ago
- a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for eas…☆20Mar 14, 2025Updated last year
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 3 months ago
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆24May 31, 2026Updated 2 months ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆17Jul 23, 2026Updated last week
- llama-benchy - llama-bench style benchmarking tool for all backends☆613Jul 10, 2026Updated 3 weeks ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 4 months ago
- A simple speech-to-text and text-to-speech AI chatbot that can be run fully offline.☆45Jan 28, 2024Updated 2 years ago
- ONNX speech pipeline library for ASR, diarization, VAD, and denoising☆19Jun 14, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆85Jul 28, 2026Updated last week
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems☆421Updated this week
- LCM OpenVINO model converter☆24Mar 27, 2024Updated 2 years ago
- ☆52Nov 17, 2025Updated 8 months ago
- Simple, Fast, Parallel Huggingface GGML model downloader written in python☆24Jul 26, 2023Updated 3 years ago
- ☆16Apr 21, 2026Updated 3 months ago
- RMON is geo-distributed monitoring of sites and infrastructure remote monitoring check monitor infrastructure web site geo monitoring☆24Jul 2, 2026Updated last month