Omni inference in C/C++
☆272Sep 27, 2026Updated this week
Alternatives and similar repositories for llama.cpp-omni
Users that are interested in llama.cpp-omni are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cook up amazing AI applications effortlessly with MiniCPM / MiniCPM-V / MiniCPM-o☆641Aug 28, 2026Updated last month
- Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5☆390Sep 15, 2026Updated last week
- Standalone C++ inference project for VoxCPM models built on top of ggml.☆94Jul 14, 2026Updated 2 months ago
- ☆22May 2, 2026Updated 4 months ago
- MiniCPM-V apps — fully offline multimodal chat on iOS / Android / HarmonyOS☆397Sep 12, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆42Jul 13, 2026Updated 2 months ago
- Port of Facebook's LLaMA model in C/C++☆113Updated this week
- TTS support with GGML☆249Oct 5, 2025Updated 11 months ago
- A framework for efficient model inference with omni-modality models☆7,084Updated this week
- ☆243Jul 18, 2026Updated 2 months ago
- A lightweight pure C++ Text-to-Speech (TTS) pipeline with OpenVINO, supporting multiple languages.☆113Sep 26, 2025Updated last year
- ☆16Feb 29, 2024Updated 2 years ago
- 在kaggle部署ChatGLM API,和ChatGPT api使用相同的调用方式☆14Jun 30, 2023Updated 3 years ago
- The SchemaPin protocol for cryptographically signing and verifying AI agent tool schemas to prevent supply-chain attacks.☆16Sep 13, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- TurboServe: Serving Streaming Video Generation Efficiently and Economically☆247Jul 12, 2026Updated 2 months ago
- a collection of skills for vllm-omni☆88Sep 21, 2026Updated last week
- On-device VAD / streaming STT / TTS / diarization in C++17 (ONNX + LiteRT) with a voice-agent pipeline. Linux, Windows, Android.☆90Sep 15, 2026Updated last week
- Tutorials of Extending and importing TVM with CMAKE Include dependency.☆16Oct 11, 2024Updated last year
- This is a real-time conversation project powered by a VoxCPM-based streaming TTS model.☆113Jul 17, 2026Updated 2 months ago
- The repository targets the OpenCL gemm function performance optimization. It compares several libraries clBLAS, clBLAST, MIOpenGemm, Inte…☆17Mar 28, 2019Updated 7 years ago
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆30Jul 4, 2026Updated 2 months ago
- OpenManus-node is a Node.js and TypeScript recreation of the mannaandpoem/OpenManus project☆29Mar 15, 2025Updated last year
- A CUDA kernel for NHWC GroupNorm for PyTorch☆26Nov 15, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is an evolving repo for the paper “From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models ”A compr…☆34Dec 23, 2025Updated 9 months ago
- 🎙️ A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!☆2,591Updated this week
- Android UI based on Skia and Yoga☆26Oct 26, 2025Updated 11 months ago
- learn TensorRT from scratch🥰☆18Sep 29, 2024Updated last year
- A self-hosted AI toolkit running locally via Docker Compose, bundling an LLM gateway, workflow automation, and a chat UI — all backed by …☆16May 17, 2026Updated 4 months ago
- a conversational finance assistant that provides users with real-time stock quotes, market news, and insights on market movers through na…☆17Apr 26, 2025Updated last year
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++☆7,422Updated this week
- This repository implements the YOLOv9 model on Jetson Orin Nano☆19Aug 28, 2024Updated 2 years ago
- A local browser reading app powered by MOSS-TTS-Nano with in-browser ONNX inference☆61May 7, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆16May 21, 2026Updated 4 months ago
- Implementation of the OmniVoice inference model from k2-fsa on Rust☆29Sep 10, 2026Updated 2 weeks ago
- Local AI text-to-speech with voice cloning and voice design, powered by GGML. C++17 port of Qwen3-TTS (QwenLM/Qwen3-TTS). 10 languages, 2…☆175Updated this week
- Implementation of Qwen3-ASR-0.6B in GGML☆114Jul 28, 2026Updated 2 months ago
- High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI☆570Sep 3, 2026Updated 3 weeks ago
- BlenderLM is a Python package that enables LLMs (Large Language Models) to control and interact with Blender, the open-source 3D creation…☆15Jun 5, 2025Updated last year
- SLIM Models by LLMWare. A streamlit app showing the capabilities for AI Agents and Function Calls.☆21Feb 11, 2024Updated 2 years ago