Tensor parallelism is all you need. Run LLMs on an AI cluster at home using any device. Distribute the workload, divide RAM usage, and increase inference speed.
☆18Nov 11, 2024Updated last year
Alternatives and similar repositories for distributed-llama
Users that are interested in distributed-llama are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 3 months ago
- Steganography Reverse Shell☆10Apr 22, 2023Updated 3 years ago
- Chat GPT Things by Taylor Newsome☆13Mar 19, 2024Updated 2 years ago
- Play with OpenAI API's using your own API Key. Your API Key is stored and used only from your browser.☆14Dec 20, 2025Updated 8 months ago
- Training HuggingFace models using fastai☆11Jul 22, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Handle Android "draw over other apps" permissions and queries in a version-agnostic way☆22Jun 17, 2019Updated 7 years ago
- Download and Transcribe X Spaces☆11Nov 16, 2024Updated last year
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- Modification of SOMPY repo with robust K-means clustering (bootstrapped SSE elbow method)☆13Apr 6, 2019Updated 7 years ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- meson android build PoC☆11Oct 29, 2019Updated 6 years ago
- Unofficial Claude Code SDKs for Typescript and Python☆15May 20, 2025Updated last year
- ☆27Updated this week
- ☆10Dec 29, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This repository is intended as a comprehensive guide to prepare for interviews focused on generative AI. It serves as a one-stop resource…☆11Dec 13, 2024Updated last year
- Flask based Web application for predicting the income of a person☆13Dec 23, 2018Updated 7 years ago
- Tools for merging pretrained large language models.☆19Jun 12, 2024Updated 2 years ago
- ☆62Aug 6, 2026Updated 3 weeks ago
- Bloom filter alternative (C++)☆18Nov 8, 2018Updated 7 years ago
- Gradient-based Hyperparameter Optimization Over Long Horizons☆14Sep 29, 2021Updated 4 years ago
- A fully offline voice assistant that combines lmstudio and applio together. Uses two methods of TTS, STT and also has some extra features…☆23Apr 14, 2025Updated last year
- Minimal, highly available (HA) Kubernetes cluster on Hetzner Cloud — up and running in under 10 minutes.☆11Apr 23, 2026Updated 4 months ago
- Specialized AI agents for your bare metal. 100% On-Prem & Air-Gap ready.☆26Feb 21, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Deploy fastai models with Docker☆19Sep 27, 2020Updated 5 years ago
- ☆16Nov 24, 2025Updated 9 months ago
- Code and data for editing model beliefs with SDF and other methods, and for evaluating the depth of the implanted beliefs.☆19Oct 23, 2025Updated 10 months ago
- A library to automate the conversion of linux-based VMs to a set of docker containers☆14Apr 10, 2015Updated 11 years ago
- ☆10Jun 24, 2026Updated 2 months ago
- Additional functionality for use with fastai’s medical imaging module☆15Jul 20, 2022Updated 4 years ago
- Low-latency ASR using SpeechBrain StreamingASR and torchaudio StreamReader.☆18Apr 19, 2025Updated last year
- Multivariate Bayesian Structural Time Series in Stan☆13Apr 13, 2020Updated 6 years ago
- ☆20May 30, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A minimal home grid world environment to evaluate language understanding in interactive agents.☆24Sep 6, 2023Updated 2 years ago
- ☆17Mar 20, 2026Updated 5 months ago
- Automate your linkedin networking using an AI agent☆19Jan 21, 2025Updated last year
- ☆13Oct 16, 2024Updated last year
- Multi-stage LLM agent pipeline for optimizing Triton kernels on Intel XPU — from analysis to autotuning.☆21Updated this week
- Kubernetes deployment strategies from "DB Schemas & Kubernetes Rollouts" blogpost☆20May 8, 2019Updated 7 years ago
- Using Llam.cpp and onnxruntime to accelerate inference of GOT-OCR2.0☆16Mar 6, 2025Updated last year