JakeATX / llamAmpereView on GitHub
llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8, supporting full 262K ctx.
☆148Oct 7, 2026Updated this week

Alternatives and similar repositories for llamAmpere

Users that are interested in llamAmpere are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?