FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors between GPU and CPU memory.
☆109Jun 3, 2026Updated last month
Alternatives and similar repositories for flextensor
Users that are interested in flextensor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆280Jul 6, 2026Updated 3 weeks ago
- Triton Model Navigator is an inference toolkit designed for optimizing and deploying Deep Learning models with a focus on NVIDIA GPUs.☆224May 27, 2026Updated 2 months ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆56Feb 25, 2026Updated 5 months ago
- PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.☆846Aug 13, 2025Updated 11 months ago
- SiMM: Scalable in-Memory Middleware☆41Apr 20, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆16May 14, 2025Updated last year
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 4 months ago
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆121Feb 28, 2026Updated 5 months ago
- Dynamo Workshop☆19Nov 7, 2025Updated 8 months ago
- Cross-GPU KV Cache Marketplace☆26Nov 12, 2025Updated 8 months ago
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 3 months ago
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI tr…☆121Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆369Jul 9, 2026Updated 3 weeks ago
- Notes and artifacts from the ONNX steering committee☆29Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆18Feb 9, 2026Updated 5 months ago
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆463Jul 14, 2026Updated 2 weeks ago
- A collection of optimizers, some arcane others well known, for Flax.☆29Aug 6, 2021Updated 4 years ago
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Jan 30, 2026Updated 5 months ago
- NVIDIA Inference Xfer Library (NIXL)☆1,157Updated this week
- Workshop materials for AI Engineer World's Fair☆18Jun 3, 2025Updated last year
- torchcomms: a modern PyTorch communications API☆381Updated this week
- Quantize transformers to any learned arbitrary 4-bit numeric format☆59Jul 2, 2026Updated 3 weeks ago
- Offline optimization of your disaggregated Dynamo graph☆379Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆423Updated this week
- KV cache store for distributed LLM inference☆423Nov 13, 2025Updated 8 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆123Apr 17, 2026Updated 3 months ago
- ☆20Feb 2, 2026Updated 5 months ago
- ☆12Mar 16, 2022Updated 4 years ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆832Jul 14, 2026Updated 2 weeks ago
- ☆310Updated this week
- ☆18Jun 2, 2026Updated last month
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆669Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Quantized LLM training in pure CUDA/C++.☆251Updated this week
- Code for "Accelerating Training with Neuron Interaction and Nowcasting Networks" [ICLR 2025]☆30Feb 20, 2026Updated 5 months ago
- mKernel: fast multi-node, multi-GPU fused kernels☆257Updated this week
- Pure C / AVX-512 port of Craftax-Classic. 47.8M SPS on a Ryzen 9 9950X3D -- 3.2x an RTX Pro 6000 Blackwell on the same env.☆24Jul 14, 2026Updated 2 weeks ago
- High-throughput tensor loading for PyTorch☆260Updated this week
- PyTorch Code for the Paper: "Exploiting Uncertainty of Loss Landscape for Stochastic Optimization [Bhaskara et al. (2019)]☆16Apr 30, 2026Updated 3 months ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆475Updated this week