☆242Sep 19, 2026Updated this week
Alternatives and similar repositories for b12x
Users that are interested in b12x are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,075Updated this week
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 3 months ago
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆29Sep 3, 2026Updated 2 weeks ago
- Test to see if a DGX Spark (or similar GB10 device) is throttling due to possible Power Delivery issues☆32Mar 9, 2026Updated 6 months ago
- LLM inference in C/C++☆30Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆509Updated this week
- An ultra-fast, distributed Safetensors loader☆78Sep 13, 2026Updated last week
- HPC Game Platform☆11Apr 20, 2023Updated 3 years ago
- Vitis 部署加速器工作流介绍☆13Jan 10, 2025Updated last year
- Manages vllm-nccl dependency☆19Jun 3, 2024Updated 2 years ago
- ☆20Jul 21, 2026Updated last month
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆103Updated this week
- Docker configuration for running VLLM on dual DGX Sparks☆2,287Updated this week
- ☆20Apr 26, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆51Aug 17, 2026Updated last month
- ☆39Dec 14, 2025Updated 9 months ago
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,254Sep 1, 2026Updated 2 weeks ago
- Agent-friendly GPU profile-query CLI☆122Aug 21, 2026Updated 3 weeks ago
- Your AI colleague, in the apps you already use.☆14Updated this week
- ☆74Feb 27, 2026Updated 6 months ago
- Pure Rust Inference Engine☆696Updated this week
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,459Updated this week
- TokenSpeed is a speed-of-light LLM inference engine.☆2,149Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Code for running forward and backward versions of GPT2☆10Nov 20, 2021Updated 4 years ago
- A benchmark framework for Tensorflow 2.X☆12Apr 20, 2024Updated 2 years ago
- easy exllama interface w/ automation & evals☆19Updated this week
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆480Sep 10, 2026Updated last week
- This is a repo covers ai research papers pseudocodes☆18Jun 20, 2023Updated 3 years ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- Unlock P2P comms between consumer NVIDIA GPUs☆539Updated this week
- 🎉My Collections of CUDA Kernels~☆11Jun 25, 2024Updated 2 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,244Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- A minimalist Docker project to help people getting started with Node, WizardCoder, CTransformers, Python, Express and TypeScript. Ready t…☆14Jun 23, 2023Updated 3 years ago
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- FLA but cuTile☆27Apr 17, 2026Updated 5 months ago
- ☆32Jul 2, 2025Updated last year
- Can LLMs Write Correct and Efficient GPU Communication Code?☆63Jul 7, 2026Updated 2 months ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆540Updated this week
- ☆16Oct 25, 2024Updated last year