Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5
☆336Aug 13, 2026Updated this week
Alternatives and similar repositories for MiniCPM-o-Demo
Users that are interested in MiniCPM-o-Demo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Omni inference in C/C++☆243Updated this week
- Cook up amazing AI applications effortlessly with MiniCPM / MiniCPM-V / MiniCPM-o☆619Updated this week
- Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.☆294Jul 17, 2026Updated 3 weeks ago
- MiniCPM-V apps — fully offline multimodal chat on iOS / Android / HarmonyOS☆364Updated this week
- DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action☆122May 20, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Fun-Audio-Chat is a Large Audio Language Model built for natural, low-latency voice interactions.☆990Feb 27, 2026Updated 5 months ago
- ☆109Oct 16, 2025Updated 9 months ago
- [ICLR 2026] MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning☆44Jan 14, 2026Updated 7 months ago
- Covo-Audio is a 7B-parameter end-to-end large audio language model that directly processes continuous audio inputs and generates audio ou…☆177Mar 17, 2026Updated 4 months ago
- Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/…☆312Updated this week
- ClawXMemory: A Multi-Level Memory Plugin for OpenClaw with Long-Term Context☆50Apr 15, 2026Updated 3 months ago
- Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.☆667Jul 27, 2026Updated 2 weeks ago
- Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation…☆1,494Mar 16, 2026Updated 4 months ago
- ProactiveBench: A Comprehensive Benchmark for VideoLLM Proactive Interaction Evaluation☆20Jan 8, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Math Formula Agent☆32Aug 6, 2026Updated last week
- Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, im…☆3,951Apr 23, 2026Updated 3 months ago
- 🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!☆2,318Aug 6, 2026Updated last week
- Official code of "RoboOmni: Proactive Robot Manipulation in Omni-modal Context"☆118Mar 28, 2026Updated 4 months ago
- This is a real-time conversation project powered by a VoxCPM-based streaming TTS model.☆87Jul 17, 2026Updated 3 weeks ago
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]☆15Jun 1, 2026Updated 2 months ago
- JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System☆1,727Updated this week
- Voca - Your local voice cloning assistant. Powered by VoxCPM☆44Aug 5, 2026Updated last week
- ☆346Apr 15, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆284Jun 17, 2026Updated last month
- ☆120Jun 5, 2026Updated 2 months ago
- ☆30Jul 17, 2026Updated 3 weeks ago
- A general purpose task-agnostic speech augmentation policy☆17Mar 13, 2026Updated 5 months ago
- ☆26Feb 10, 2026Updated 6 months ago
- The implementation of "End-to-End Neural Speaker Diarization with an Iterative Adaptive Attractor Estimation", which is accepted by Neura…☆11Aug 27, 2023Updated 2 years ago
- A Benchmark for Evaluating Turn-Taking and Overlap Handling in Full-Duplex Spoken Dialogue Models☆263May 20, 2026Updated 2 months ago
- 这是一个面向家庭场景的离线监控视频分析系统。项目以 NAS/目录中的摄像头录像为输入,自动完成视频扫描、会话合并、AI 事件识别、家庭日报生成,并提供 Web 管理后台、自然语言问答、Webhook 与 MCP 对外能力。☆21Jun 14, 2026Updated 2 months ago
- ☆291Aug 7, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A curated list of full-duplex spoken dialogue models & benchmarks☆182Updated this week
- ☆693Apr 29, 2026Updated 3 months ago
- ☆33Aug 6, 2026Updated last week
- MiMo-Audio: Audio Language Models are Few-Shot Learners☆1,076Jun 17, 2026Updated last month
- [TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants☆383Updated this week
- ☆110Aug 7, 2026Updated last week
- CoDeTT测试集代码、数据集链接以及论文地址☆16Jun 6, 2026Updated 2 months ago