test1111111111111112 / llama-cpp-turboquant-gemma4View on GitHub
TurboQuant llama.cpp fork with optimized turbo4 kernels for Gemma 4 D=256/512 heads — lazy K/V, batch decode, warp-cooperative write. 120 t/s with 3.8x KV compression on RTX 3090.
35Apr 5, 2026Updated 3 months ago

Alternatives and similar repositories for llama-cpp-turboquant-gemma4

Users that are interested in llama-cpp-turboquant-gemma4 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?