Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
☆486Jul 16, 2026Updated this week
Alternatives and similar repositories for OpenArc
Users that are interested in OpenArc are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆429Updated this week
- Make use of Intel Arc Series GPU to Run Ollama, StableDiffusion, Whisper and Open WebUI, for image generation, speech recognition and int…☆382Jul 14, 2026Updated last week
- AI PC starter app for doing AI image creation, image stylizing, and chatbot on a PC powered by an Intel® Arc™ GPU.☆937Updated this week
- ☆54Updated this week
- Add genai backend for ollama to run generative AI models using OpenVINO Runtime.☆30Apr 16, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The vLLM XPU kernels for Intel GPU☆55Updated this week
- ☆180Updated this week
- An fully autonomous agent that accesses the browser and performs tasks.☆18Apr 25, 2025Updated last year
- Run Generative AI models with simple C++/Python API and using OpenVINO Runtime☆556Updated this week
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- This repository contains Dockerfiles, scripts, yaml files, Helm charts, etc. used to scale out AI containers with versions of TensorFlow …☆79May 27, 2026Updated last month
- OpenVINO Intel NPU Compiler☆92Updated this week
- Surgically de-slop LLMs☆15Jun 1, 2025Updated last year
- ☆97Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- The AI PC Application Installer provides a unified way to set up Intel AI PC development environments.☆40May 21, 2026Updated 2 months ago
- Using OpenVINO to speed up MeloTTS inference☆15Nov 1, 2024Updated last year
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,066Updated this week
- Local-first desktop AI workbench for roleplay, multi-character chat, long-form writing, RAG, MCP tools, plugins, and local models.☆104Updated this week
- A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD),…☆21Updated this week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆34Apr 20, 2026Updated 3 months ago
- Run multiple resource-heavy Large Models (LM) on the same machine with limited amount of VRAM/other resources by exposing them on differe…☆89Jul 13, 2026Updated last week
- Implementation of Sesame's Conversational Speech Model for Hugging Face Transformers☆58May 17, 2025Updated last year
- Privacy-first agentic framework with powerful reasoning & task automation capabilities. Natively distributed and fully ISO 27XXX complian…☆70Apr 1, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An inference engine built with SYCL + oneDNN☆22Updated this week
- ☆19Aug 19, 2025Updated 11 months ago
- 🤗 Optimum Intel: Accelerate inference with Intel optimization tools☆606Updated this week
- Fully automated installation scripts for ComfyUI optimized for Intel Arc GPUs (A-Series) and Intel Core Ultra iGPUs with XPU backend, Tri…☆154Feb 10, 2026Updated 5 months ago
- A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support…☆1,530Updated this week
- ☆99Mar 28, 2026Updated 3 months ago
- Adapt IPEX to CUDA☆46Jul 14, 2026Updated last week
- llama.cpp fork with additional SOTA quants and improved performance☆2,943Updated this week
- Fused BF16 Huffman GEMV Inference kernel☆22Apr 22, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆24Apr 1, 2025Updated last year
- Local runner for Microsoft VibeVoice Realtime TTS Fully compatible with Open-Webui Plug and Play. OpenAI api endpoint .Run the Colab note…☆43Updated this week
- OpenVINO Tokenizers extension☆53Updated this week
- ☆261Jun 4, 2025Updated last year
- Tools for easier OpenVINO development/debugging☆10Jul 16, 2025Updated last year
- ☆202Mar 31, 2025Updated last year
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆4,997Updated this week