sergiuszm / ninfer-4090View on GitHub
Qwen3.8-27B on one RTX 4090: full native 262K context via E8 4-bit KV, up to 149 tok/s code decode with MTP3, sm_89-retuned attention prefill, vision, llama.cpp-compatible /metrics + /slots
136Sep 12, 2026Updated this week

Alternatives and similar repositories for ninfer-4090

Users that are interested in ninfer-4090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?