☆194Aug 28, 2026Updated this week
Alternatives and similar repositories for b12x
Users that are interested in b12x are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆79Updated this week
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆25Apr 28, 2026Updated 4 months ago
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆17Jun 5, 2026Updated 2 months ago
- DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)☆26Aug 6, 2026Updated 3 weeks ago
- Test to see if a DGX Spark (or similar GB10 device) is throttling due to possible Power Delivery issues☆31Mar 9, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- AEON vLLM Ultimate — vLLM 0.27.1 built from source for DGX Spark / Blackwell (sm_121a/GB10). DSpark quantized Markov heads, DFlash SWA on…☆138Updated this week
- LLM inference in C/C++☆29Updated this week
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆483Updated this week
- An ultra-fast, distributed Safetensors loader☆72Updated this week
- Vitis 部署加速器工作流介绍☆13Jan 10, 2025Updated last year
- Manages vllm-nccl dependency☆19Jun 3, 2024Updated 2 years ago
- ☆20Jul 21, 2026Updated last month
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆103Updated this week
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆21Aug 17, 2026Updated 2 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆20Apr 26, 2026Updated 4 months ago
- ☆51Aug 17, 2026Updated 2 weeks ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,202Updated this week
- ☆16Jan 15, 2026Updated 7 months ago
- Agent-friendly GPU profile-query CLI☆119Aug 21, 2026Updated last week
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,238Jul 30, 2026Updated last month
- ☆73Feb 27, 2026Updated 6 months ago
- Pure Rust Inference Engine☆672Updated this week
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,201Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- TokenSpeed is a speed-of-light LLM inference engine.☆2,042Updated this week
- A benchmark framework for Tensorflow 2.X☆12Apr 20, 2024Updated 2 years ago
- easy exllama interface w/ automation & evals☆17Aug 11, 2026Updated 2 weeks ago
- DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding☆54Jun 28, 2026Updated 2 months ago
- NVIDIA Linux open GPU with P2P support☆459Aug 22, 2026Updated last week
- 🎉My Collections of CUDA Kernels~☆11Jun 25, 2024Updated 2 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- FLA but cuTile☆27Apr 17, 2026Updated 4 months ago
- ☆32Jul 2, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A high-throughput and memory-efficient inference and serving engine for LLMs☆41Aug 13, 2026Updated 2 weeks ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆536Updated this week
- A hands-on guide for AI builders: make your own RTX 4090D/5090 GPU server that’s fast and efficient.☆17Aug 19, 2025Updated last year
- Can LLMs Write Correct and Efficient GPU Communication Code?☆63Jul 7, 2026Updated last month
- ☆775Aug 23, 2026Updated last week
- Benchmarking tool for vLLM inference performance with GPU monitoring☆52Jun 7, 2026Updated 2 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆155Aug 21, 2026Updated last week