b4rtaz / distributed-llama

Tensor parallelism is all you need. Run LLMs on an AI cluster at home using any device. Distribute the workload, divide RAM usage, and increase inference speed.
β˜†1,690Updated this week

Alternatives and similar repositories for distributed-llama:

Users that are interested in distributed-llama are comparing it to the libraries listed below