High-performance safetensors model loader
☆167Sep 10, 2026Updated last week
Alternatives and similar repositories for fastsafetensors
Users that are interested in fastsafetensors are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An ultra-fast, distributed Safetensors loader☆78Sep 13, 2026Updated last week
- "An optimizer custom node for ComfyUI that ensures each queue execution starts in an optimal state by clearing unused VRAM and unnecessar…☆23Jul 18, 2025Updated last year
- ☆382Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,263Updated this week
- KV cache store for distributed LLM inference☆439Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A model loader that uses fastsafetensors library to perform a fast, zero-copy load from storage to VRAM.☆23Jan 21, 2026Updated 7 months ago
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆156Updated this week
- Fast and memory-efficient exact attention☆23Jun 26, 2026Updated 2 months ago
- (ECCV 2026): Official code for Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models☆23Jul 9, 2026Updated 2 months ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆49May 13, 2025Updated last year
- DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling☆43Updated this week
- htop-like TUI for real-time RDMA network monitoring.☆104Updated this week
- Module, Model, and Tensor Serialization/Deserialization☆326Jul 7, 2026Updated 2 months ago
- A lightweight design for computation-communication overlap.☆247Jan 20, 2026Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Offline optimization of your disaggregated Dynamo graph☆446Updated this week
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆31Jan 22, 2026Updated 7 months ago
- A NCCL extension library, designed to efficiently offload GPU memory allocated by the NCCL communication library.☆119Dec 17, 2025Updated 9 months ago
- Modular RDMA Interface☆180Updated this week
- VUA stands for 'VAST Undivided Attention'. It's a global KVCache storage solution optimizing LLM time to first token (TTFT) and GPU utili…☆38Mar 12, 2026Updated 6 months ago
- NVIDIA GPUDirect Storage Driver☆384Updated this week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.