FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors between GPU and CPU memory.
☆111Aug 18, 2026Updated 3 weeks ago
Alternatives and similar repositories for flextensor
Users that are interested in flextensor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆281Updated this week
- Triton Model Navigator is an inference toolkit designed for optimizing and deploying Deep Learning models with a focus on NVIDIA GPUs.☆225May 27, 2026Updated 3 months ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆61Feb 25, 2026Updated 6 months ago
- PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.☆848Aug 13, 2025Updated last year
- SiMM: Scalable in-Memory Middleware☆42Apr 20, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆16May 14, 2025Updated last year
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆123Feb 28, 2026Updated 6 months ago
- Dynamo Workshop☆19Nov 7, 2025Updated 10 months ago
- Cross-GPU KV Cache Marketplace☆27Nov 12, 2025Updated 9 months ago
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 4 months ago
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI tr…☆152Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆381Jul 9, 2026Updated last month
- Notes and artifacts from the ONNX steering committee☆29Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆19Feb 9, 2026Updated 6 months ago
- A collection of optimizers, some arcane others well known, for Flax.☆29Aug 6, 2021Updated 5 years ago
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆479Aug 31, 2026Updated last week
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Jan 30, 2026Updated 7 months ago
- NVIDIA Inference Xfer Library (NIXL)☆1,235Updated this week
- Quantize transformers to any learned arbitrary 4-bit numeric format☆59Jul 2, 2026Updated 2 months ago
- torchcomms: a modern PyTorch communications API☆391Updated this week
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆144Updated this week
- Offline optimization of your disaggregated Dynamo graph☆437Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆511Updated this week
- Workshop materials for AI Engineer World's Fair☆21Jun 3, 2025Updated last year
- KV cache store for distributed LLM inference☆436Nov 13, 2025Updated 9 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆129Apr 17, 2026Updated 4 months ago
- ☆20Feb 2, 2026Updated 7 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆847Aug 4, 2026Updated last month
- ☆338Updated this week
- ☆18Jun 2, 2026Updated 3 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆812Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Quantized LLM training in pure CUDA/C++.☆257Aug 31, 2026Updated last week
- Code for "Accelerating Training with Neuron Interaction and Nowcasting Networks" [ICLR 2025]☆30Feb 20, 2026Updated 6 months ago
- mKernel: fast multi-node, multi-GPU fused kernels☆270Updated this week
- Pure C / AVX-512 port of Craftax-Classic. 47.8M SPS on a Ryzen 9 9950X3D -- 3.2x an RTX Pro 6000 Blackwell on the same env.☆24Jul 14, 2026Updated last month
- High-throughput tensor loading for PyTorch☆266Aug 17, 2026Updated 3 weeks ago
- PyTorch Code for the Paper: "Exploiting Uncertainty of Loss Landscape for Stochastic Optimization [Bhaskara et al. (2019)]☆16Apr 30, 2026Updated 4 months ago
- AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solu…☆646Updated this week