Use safetensors with ONNX 🤗
☆88Jun 23, 2026Updated 3 weeks ago
Alternatives and similar repositories for onnx-safetensors
Users that are interested in onnx-safetensors are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Some benchmarks☆12Sep 19, 2019Updated 6 years ago
- MXNet bindings for .NET/F#☆15May 26, 2021Updated 5 years ago
- The backend behind the LLM-Perf Leaderboard☆11May 5, 2024Updated 2 years ago
- Visualize ONNX models with model-explorer☆74Jun 30, 2026Updated 3 weeks ago
- ANE accelerated embedding models!☆20Dec 11, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ONNX format parsing and manipulation in C#.☆36Jan 16, 2025Updated last year
- 🤗 Optimum ONNX: Export your model to ONNX and run inference with ONNX Runtime☆153Updated this week
- Model compression for ONNX☆102May 1, 2026Updated 2 months ago
- QONNX: Arbitrary-Precision Quantized Neural Networks in ONNX☆191Jun 10, 2026Updated last month
- A lightweight, single-header C++11 Jinja2 template engine for LLM chat templates.☆20Mar 4, 2026Updated 4 months ago
- ☆42Nov 29, 2022Updated 3 years ago
- Large Language Model Onnx Inference Framework☆35Nov 25, 2025Updated 7 months ago
- Profile your CoreML models directly from Python 🐍☆29Sep 8, 2025Updated 10 months ago
- Experimental wasm32-unknown-wasi runtime for Python code execution☆40Nov 28, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- the python api for axengine runtime☆26Mar 24, 2026Updated 3 months ago
- Hugging Face Inference Toolkit used to serve transformers, sentence-transformers, and diffusers models.☆94May 28, 2026Updated last month
- Find out why your CoreML model isn't running on the Neural Engine!☆30Jun 18, 2024Updated 2 years ago
- The Triton backend for the ONNX Runtime.☆183Updated this week
- ONNX Script enables developers to naturally author ONNX functions and models using a subset of Python.☆447Updated this week
- ☆31May 10, 2026Updated 2 months ago
- Tutorial on how to convert machine learned models into ONNX☆14Mar 11, 2023Updated 3 years ago
- A general 2-8 bits quantization toolbox with GPTQ/AWQ/HQQ/VPTQ, and export to onnx/onnx-runtime easily.☆190Mar 23, 2026Updated 3 months ago
- A Python wrapper around HuggingFace's TGI (text-generation-inference) and TEI (text-embedding-inference) servers.☆32Sep 19, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆23Jan 3, 2024Updated 2 years ago
- A lightweight, production-ready C++ library for LLM tokenization, fully compatible with HuggingFace tokenizer.json.☆33Jan 4, 2026Updated 6 months ago
- Comfyui implementation of Regional Adaptive Sampling, for Flux and HunYuanVideo☆23Jun 30, 2026Updated 3 weeks ago
- Facial recognition engine☆10Jul 12, 2026Updated last week
- onnxruntime-extensions: A specialized pre- and post- processing library for ONNX Runtime☆474Updated this week
- Extension for Automatic1111's Stable Diffusion WebUI, using Microsoft DirectML to deliver high performance result on any Windows GPU.☆65Jun 5, 2024Updated 2 years ago
- ☆13Feb 1, 2022Updated 4 years ago
- ☆16Apr 23, 2024Updated 2 years ago
- Convert KataGo network files to ONNX format.☆15Nov 7, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- No-code CLI designed for accelerating ONNX workflows☆241Jul 1, 2026Updated 2 weeks ago
- Multi-stream video inference with Ultralytics YOLO - Display multiple video streams in a grid layout with real-time object detection.☆16May 20, 2026Updated 2 months ago
- 8-bit floating point types for Rust☆64Feb 4, 2026Updated 5 months ago
- example of using CoreML from c++☆24Jun 14, 2023Updated 3 years ago
- ONNX Optimizer☆825Updated this week
- Express.js ported to a Service Worker context☆18Mar 6, 2025Updated last year
- ☆16Aug 10, 2022Updated 3 years ago