TPU inference for vLLM, with unified JAX and PyTorch support.
☆391Jul 26, 2026Updated this week
Alternatives and similar repositories for tpu-inference
Users that are interested in tpu-inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- JAX backend for SGL☆315Updated this week
- Tokamax: A GPU and TPU kernel library.☆254Updated this week
- ☆119Updated this week
- torchax is a PyTorch frontend for JAX. It gives JAX the ability to author JAX programs using familiar PyTorch syntax. It also provides JA…☆232Jul 3, 2026Updated 3 weeks ago
- Minimal yet performant LLM examples in pure JAX☆269Jul 4, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆47Updated this week
- ☆377Updated this week
- Microbenchmarking hyperparameter tuning for JAX functions.☆22Jul 7, 2026Updated 2 weeks ago
- ☆22Jul 7, 2026Updated 2 weeks ago
- A Lightweight LLM Post-Training Library☆2,386Updated this week
- JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs wel…☆451Jan 5, 2026Updated 6 months ago
- A simple, performant, and scalable Jax LLM!☆2,366Updated this week
- a Jax quantization library☆126Updated this week
- ☆35Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A simplified and automated orchestration workflow to perform ML end-to-end (E2E) model tests and benchmarking on Cloud VMs across differe…☆64Updated this week
- Orbax provides common checkpointing and persistence utilities for JAX users☆526Updated this week
- Recipes for reproducing training and serving benchmarks for large machine learning models using GPUs on Google Cloud.☆138Updated this week
- ☆591Jul 11, 2024Updated 2 years ago
- MLIR-based partitioning system☆198Updated this week
- xpk (Accelerated Processing Kit, pronounced x-p-k,) is a software tool to help Cloud developers to orchestrate training jobs on accelerat…☆193Updated this week
- PyTorch/XLA integration with JetStream (https://github.com/google/JetStream) for LLM inference"☆86Dec 18, 2025Updated 7 months ago
- easydel jax kernels writen in triton for gpus and pallas for tpus☆28Updated this week
- Minimal, lightweight JAX implementations of popular models.☆239May 29, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Google TPU optimizations for transformers models☆135Jan 23, 2026Updated 6 months ago
- Package of Pathways-on-Cloud utilities☆31Updated this week
- ☆15May 11, 2025Updated last year
- Automatic differentiation for Triton Kernels☆29Aug 12, 2025Updated 11 months ago
- Testing framework for Deep Learning models (Tensorflow and PyTorch) on Google Cloud hardware accelerators (TPU and GPU)☆64Updated this week
- Pax is a Jax-based machine learning framework for training large scale models. Pax allows for advanced and fully configurable experimenta…☆555Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆912Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B2…☆1,282Updated this week
- ☆32Jul 9, 2026Updated 2 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,391Updated this week
- PyTorch Single Controller☆1,065Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,037Updated this week
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆36Mar 20, 2026Updated 4 months ago
- JaxPP is a library for JAX that enables flexible MPMD pipeline parallelism for large-scale LLM training☆82Jun 18, 2026Updated last month
- NVIDIA Inference Xfer Library (NIXL)☆1,153Updated this week
- SGLang kernel library for Intel XPU☆27Updated this week