Accelerate LLM with low-bit (FP4 / INT4 / FP8 / INT8) optimizations using ipex-llm
☆171May 4, 2026Updated 4 months ago
Alternatives and similar repositories for ipex-llm-tutorial
Users that are interested in ipex-llm-tutorial are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V,…☆8,857Jan 28, 2026Updated 7 months ago
- Step-by-step Deep Leaning Tutorials on Apache Spark using BigDL☆210Jan 3, 2023Updated 3 years ago
- 🤗 Optimum Intel: Accelerate inference with Intel optimization tools☆617Updated this week
- A Gradio Web UI for running local LLM on Intel GPU (e.g., local PC with iGPU, discrete GPU such as Arc, Flex and Max) using IPEX-LLM.☆17Aug 30, 2026Updated last week
- This is a personal archive. Please refer to github.com/UCLA-VAST/RapidStream☆15May 31, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NAACL 2025] MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning☆23May 31, 2025Updated last year
- An OpenAI API compatible images server to generate or manipulate images.☆18Feb 2, 2025Updated last year
- IBM Quantum Challenge Fall 2023☆10May 23, 2023Updated 3 years ago
- Building reliable Retrieval Augmented Generation(RAG) AI Architecture☆13Jul 30, 2024Updated 2 years ago
- xeCJK使用范例说明解析☆14Feb 27, 2020Updated 6 years ago
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- Demo on iGPU for FFmpeg decode and scale, OpenVINO inference. this is zero-copy solution, which means No frame data copy from CPU to iGPU…☆17Jan 25, 2023Updated 3 years ago
- In this course navigates through the LLMOps pipeline, enabling you to preprocess training data for supervised fine-tuning and deploy cust…☆15Feb 13, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆14Apr 22, 2024Updated 2 years ago
- ☆20Feb 18, 2025Updated last year
- Projects completed under LinuxWorld Informatics Ltd. - MLOps Training.☆12Aug 15, 2020Updated 6 years ago
- Intel® Tensor Processing Primitives extension for Pytorch*☆19Aug 24, 2026Updated 2 weeks ago
- 基于 ZeroMQ 封装的进程间通信库,支持按 Topic 过滤的发布订阅模式和 RPC 模式通信☆15Feb 2, 2023Updated 3 years ago
- Turn PostgreSQL into your search engine in a Pythonic way.☆52Aug 29, 2025Updated last year
- A simple no-install web UI for Ollama and OAI-Compatible APIs!☆31Jan 30, 2025Updated last year
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Official implementation "ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations"☆21Oct 29, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Synthetic data for fine tuning LLM☆28Dec 26, 2024Updated last year
- Website and protocol specification☆14Feb 3, 2025Updated last year
- Workshop for Model Context Protocol☆17Mar 27, 2025Updated last year
- springboot2 springsecurity gradle 整合☆14May 19, 2020Updated 6 years ago
- Graphical user interface for tensor networks☆13Jul 27, 2020Updated 6 years ago
- An open-sourced PyTorch library for developing energy efficient multiplication-less models and applications.☆14Feb 3, 2025Updated last year
- ☆16May 27, 2026Updated 3 months ago
- A scalable inference server for models optimized with OpenVINO™☆926Updated this week
- MicroSIP softphone auto configuration script for using in Enterprise☆21Jun 19, 2022Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- AUTOSAR MCAL 4.4.0☆16Apr 29, 2026Updated 4 months ago
- The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.☆14Mar 30, 2024Updated 2 years ago
- Limit Orderbook Replay/Analysis Library☆10Nov 19, 2018Updated 7 years ago
- Pipeline used internally for Peter Bubenik's TDA Group at UF.☆11Nov 3, 2022Updated 3 years ago
- ☆14Mar 30, 2026Updated 5 months ago
- VS Code workspace template for app and image developers☆15May 21, 2026Updated 3 months ago
- L1-based dictionary learning for sparse coding☆11Oct 9, 2024Updated last year