FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors between GPU and CPU memory.
☆110Aug 3, 2026Updated 2 weeks ago
Alternatives and similar repositories for flextensor
Users that are interested in flextensor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆280Aug 4, 2026Updated 2 weeks ago
- Triton Model Navigator is an inference toolkit designed for optimizing and deploying Deep Learning models with a focus on NVIDIA GPUs.☆225May 27, 2026Updated 2 months ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆57Feb 25, 2026Updated 5 months ago
- SiMM: Scalable in-Memory Middleware☆42Apr 20, 2026Updated 3 months ago
- ☆16May 14, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆120Feb 28, 2026Updated 5 months ago
- Dynamo Workshop☆19Nov 7, 2025Updated 9 months ago
- Cross-GPU KV Cache Marketplace☆27Nov 12, 2025Updated 9 months ago
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 4 months ago
- vLLM plugin for HyperCLOVAX☆15Jan 27, 2026Updated 6 months ago
- ☆28Nov 6, 2024Updated last year
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI tr…☆147Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆374Jul 9, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Notes and artifacts from the ONNX steering committee☆29Updated this week
- ☆18Feb 9, 2026Updated 6 months ago
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆467Jul 14, 2026Updated last month
- A collection of optimizers, some arcane others well known, for Flax.☆29Aug 6, 2021Updated 5 years ago
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Jan 30, 2026Updated 6 months ago
- Mini AI Developer☆20Mar 17, 2026Updated 5 months ago
- NVIDIA Inference Xfer Library (NIXL)☆1,195Updated this week
- Workshop materials for AI Engineer World's Fair☆19Jun 3, 2025Updated last year
- Quantize transformers to any learned arbitrary 4-bit numeric format☆59Jul 2, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- torchcomms: a modern PyTorch communications API☆390Updated this week
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆125Updated this week
- Offline optimization of your disaggregated Dynamo graph☆412Updated this week
- KV cache store for distributed LLM inference☆430Nov 13, 2025Updated 9 months ago
- ☆23Jun 8, 2021Updated 5 years ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆127Apr 17, 2026Updated 4 months ago
- ☆20Feb 2, 2026Updated 6 months ago
- ☆12Mar 16, 2022Updated 4 years ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆842Aug 4, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆320Updated this week
- Educational reference: NVIDIA Blackwell SM100 vs SM120, NVFP4, tcgen05, MoE inference on consumer Blackwell☆22Apr 28, 2026Updated 3 months ago
- ☆18Jun 2, 2026Updated 2 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆739Updated this week
- Quantized LLM training in pure CUDA/C++.☆257Aug 6, 2026Updated last week
- Code for "Accelerating Training with Neuron Interaction and Nowcasting Networks" [ICLR 2025]☆30Feb 20, 2026Updated 5 months ago
- mKernel: fast multi-node, multi-GPU fused kernels☆266Updated this week