poisonxa16 / pxq_llama.cppView on GitHub
PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16GB card. Fork of ik_llama.cpp.
12Aug 1, 2026Updated this week

Alternatives and similar repositories for pxq_llama.cpp

Users that are interested in pxq_llama.cpp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?