π€ Optimum ONNX: Export your model to ONNX and run inference with ONNX Runtime
β159Jul 21, 2026Updated 2 weeks ago
Alternatives and similar repositories for optimum-onnx
Users that are interested in optimum-onnx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π€ Optimum ExecuTorchβ134Jul 22, 2026Updated last week
- A Python wrapper around HuggingFace's TGI (text-generation-inference) and TEI (text-embedding-inference) servers.β32Sep 19, 2025Updated 10 months ago
- π€ Tokenizers.js: A pure JS/TS implementation of today's most used tokenizersβ53Jul 28, 2026Updated last week
- Use safetensors with ONNX π€β89Jul 21, 2026Updated 2 weeks ago
- π· Build compute kernelsβ213Apr 6, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ποΈ A unified multi-backend utility for benchmarking Transformers, Timm, PEFT, Diffusers and Sentence-Transformers with full support of Oβ¦β338May 26, 2026Updated 2 months ago
- Repository for ONNX SIG artifactsβ26Updated this week
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,455Updated this week
- Large Language Model Onnx Inference Frameworkβ35Nov 25, 2025Updated 8 months ago
- Knowledgeable Embedding: Injecting dynamically updatable entity knowledge into embeddings to enhance RAGβ15Aug 31, 2025Updated 11 months ago
- A Toolkit to Help Optimize Onnx Modelβ507Updated this week
- Generative AI extensions for onnxruntimeβ1,096Updated this week
- Tutorial on how to convert machine learned models into ONNXβ14Mar 11, 2023Updated 3 years ago
- Pythonic framework for building ONNX graphsβ99Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Smart commit messagesβ18Oct 25, 2024Updated last year
- β126Dec 15, 2023Updated 2 years ago
- A curated list for Efficient Large Language Modelsβ11Mar 25, 2024Updated 2 years ago
- Efficient in-memory representation for ONNX, in Pythonβ45Updated this week
- π Fine-tune OpenAI models for text classification, question answering, and moreβ17May 1, 2023Updated 3 years ago
- Google TPU optimizations for transformers modelsβ135Jan 23, 2026Updated 6 months ago
- A DDR3 Controller that uses the Xilinx MIG-7 PHY to interface with DDR3 devices.β12Aug 22, 2021Updated 4 years ago
- Benchmark tests supporting the TiledCUDA library.β19Nov 19, 2024Updated last year
- mnn tts demo.β19May 7, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Common utilities for ONNX convertersβ305Dec 16, 2025Updated 7 months ago
- β78May 14, 2026Updated 2 months ago
- A lightweight, single-header C++11 Jinja2 template engine for LLM chat templates.β20Mar 4, 2026Updated 5 months ago
- Build compute kernels and load them from the Hub.β720Updated this week
- Code for the examples presented in the talk "Training a Llama in your backyard: fine-tuning very large models on consumer hardware" givenβ¦β15Oct 16, 2023Updated 2 years ago
- A pytorch quantization backend for optimumβ1,051Updated this week
- Multiple GEMM operators are constructed with cutlass to support LLM inference.β20Aug 3, 2025Updated last year
- Implementation of the dilated self attention as described in "LongNet: Scaling Transformers to 1,000,000,000 Tokens"β13Jul 23, 2023Updated 3 years ago
- onnxruntime-extensions: A specialized pre- and post- processing library for ONNX Runtimeβ474Updated this week
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- onnxruntime-qnn is the Qualcomm AI Runtime (QAIRT) execution provider for onnxruntime. It provides onnxruntime hardware acceleration and β¦β42Updated this week
- mnn asr demo.β27Mar 24, 2025Updated last year
- A RAG that can scale π§π»βπ»β11May 28, 2024Updated 2 years ago
- Awesome code, projects, books, etc. related to CUDAβ38Jun 2, 2026Updated 2 months ago
- Qwen-Image's DiT inference with TensorRT-10β21Oct 13, 2025Updated 9 months ago
- tsingmicro AI model zooβ12Aug 6, 2025Updated 11 months ago
- Several optimization methods of half-precision general matrix vector multiplication (HGEMV) using CUDA core.β75Sep 8, 2024Updated last year