FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors between GPU and CPU memory.
☆115Aug 18, 2026Updated last month
Alternatives and similar repositories for flextensor
Users that are interested in flextensor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.☆282Updated this week
- Triton Model Navigator is an inference toolkit designed for optimizing and deploying Deep Learning models with a focus on NVIDIA GPUs.☆226May 27, 2026Updated 4 months ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆63Feb 25, 2026Updated 7 months ago
- PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.☆849Aug 13, 2025Updated last year
- SiMM: Scalable in-Memory Middleware☆43Sep 24, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆17May 14, 2025Updated last year
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Sep 18, 2026Updated 2 weeks ago
- moved to https://github.com/imjasonh/playground/tree/main/pasta☆15Aug 13, 2026Updated last month
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆124Feb 28, 2026Updated 7 months ago
- Dynamo Workshop☆19Nov 7, 2025Updated 10 months ago
- Example universal Javascript app with no Webpack/Babel using Snowpack☆13Jan 7, 2023Updated 3 years ago
- Cross-GPU KV Cache Marketplace☆27Nov 12, 2025Updated 10 months ago
- A full port of the JS/rust automerge codebase to pure-go☆21Mar 26, 2026Updated 6 months ago
- Agent application/benchmark/workload traces should be placed here.☆17Apr 13, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- vLLM plugin for HyperCLOVAX☆15Jan 27, 2026Updated 8 months ago
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI tr…☆154Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆383Jul 9, 2026Updated 2 months ago
- Notes and artifacts from the ONNX steering committee☆29Sep 24, 2026Updated last week
- ☆19Feb 9, 2026Updated 7 months ago
- A collection of optimizers, some arcane others well known, for Flax.☆29Aug 6, 2021Updated 5 years ago
- ArcticInference: vLLM plugin for high-throughput, low-latency inference☆487Sep 23, 2026Updated last week
- 🚀 Collection of libraries used with fms-hf-tuning to accelerate fine-tuning and training of large models.☆14Sep 8, 2026Updated 3 weeks ago
- NVIDIA Inference Xfer Library (NIXL)☆1,283Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Quantize transformers to any learned arbitrary 4-bit numeric format☆59Jul 2, 2026Updated 3 months ago
- torchcomms: a modern PyTorch communications API☆399Updated this week
- Offline optimization of your disaggregated Dynamo graph☆454Sep 18, 2026Updated 2 weeks ago
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆177Updated this week
- KV cache store for distributed LLM inference☆445Sep 16, 2026Updated 2 weeks ago
- Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini &…☆553Updated this week
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆131Apr 17, 2026Updated 5 months ago
- ☆20Feb 2, 2026Updated 8 months ago
- ☆12Mar 16, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆854Aug 4, 2026Updated last month
- ☆360Updated this week
- ☆18Jun 2, 2026Updated 4 months ago
- Quantized LLM training in pure CUDA/C++.☆261Aug 31, 2026Updated last month
- Code for "Accelerating Training with Neuron Interaction and Nowcasting Networks" [ICLR 2025]☆31Feb 20, 2026Updated 7 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆860Updated this week
- Pure C / AVX-512 port of Craftax-Classic. 47.8M SPS on a Ryzen 9 9950X3D -- 3.2x an RTX Pro 6000 Blackwell on the same env.☆24Jul 14, 2026Updated 2 months ago