Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5
☆389Sep 15, 2026Updated last week
Alternatives and similar repositories for MiniCPM-o-Demo
Users that are interested in MiniCPM-o-Demo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Omni inference in C/C++☆269Updated this week
- Cook up amazing AI applications effortlessly with MiniCPM / MiniCPM-V / MiniCPM-o☆639Aug 28, 2026Updated 3 weeks ago
- Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.☆318Jul 17, 2026Updated 2 months ago
- MiniCPM-V apps — fully offline multimodal chat on iOS / Android / HarmonyOS☆396Sep 12, 2026Updated last week
- DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action☆141May 20, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Fun-Audio-Chat is a Large Audio Language Model built for natural, low-latency voice interactions.☆1,009Updated this week
- ☆111Oct 16, 2025Updated 11 months ago
- [ICLR 2026] MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning☆45Jan 14, 2026Updated 8 months ago
- Covo-Audio is a 7B-parameter end-to-end large audio language model that directly processes continuous audio inputs and generates audio ou…☆182Mar 17, 2026Updated 6 months ago
- Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/…☆325Sep 16, 2026Updated last week
- Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.☆671Jul 27, 2026Updated last month
- Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation…☆1,520Mar 16, 2026Updated 6 months ago
- ProactiveBench: A Comprehensive Benchmark for VideoLLM Proactive Interaction Evaluation☆21Jan 8, 2026Updated 8 months ago
- Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, im…☆4,029Apr 23, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 🎙️ A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!☆2,580Updated this week
- Math Formula Agent☆36Updated this week
- A proactive robot manipulation model for multimodal physical environments☆120Mar 28, 2026Updated 5 months ago
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]☆15Jun 1, 2026Updated 3 months ago
- This is a real-time conversation project powered by a VoxCPM-based streaming TTS model.☆112Jul 17, 2026Updated 2 months ago
- ☆454Apr 15, 2026Updated 5 months ago
- JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System☆1,934Sep 15, 2026Updated last week
- ☆301Aug 20, 2026Updated last month
- ☆127Jun 5, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆32Aug 18, 2026Updated last month
- A general purpose task-agnostic speech augmentation policy☆17Mar 13, 2026Updated 6 months ago
- ☆27Feb 10, 2026Updated 7 months ago
- The implementation of "End-to-End Neural Speaker Diarization with an Iterative Adaptive Attractor Estimation", which is accepted by Neura…☆11Aug 27, 2023Updated 3 years ago
- 这是一个面向家庭场景的离线监控视频分析系统。项目以 NAS/目录中的摄像头录像为输入,自动完成视频扫描、会话合并、AI 事件识别、家庭日报生成,并提供 Web 管理后台、自然语言问答、Webhook 与 MCP 对外能力。☆24Sep 9, 2026Updated 2 weeks ago
- A Benchmark for Evaluating Turn-Taking and Overlap Handling in Full-Duplex Spoken Dialogue Models☆301May 20, 2026Updated 4 months ago
- ☆35Updated this week
- ☆699Apr 29, 2026Updated 4 months ago
- MiMo-Audio: Audio Language Models are Few-Shot Learners☆1,083Jun 17, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants☆396Sep 10, 2026Updated 2 weeks ago
- ☆124Aug 7, 2026Updated last month
- A curated list of full-duplex spoken dialogue models & benchmarks☆248Updated this week
- A Fully Self-Hosted Solution for Full-Duplex Voice Interaction☆589Sep 28, 2025Updated 11 months ago
- CoDeTT测试集代码、数据集链接以及论文地址☆15Jun 6, 2026Updated 3 months ago
- ☆145Feb 17, 2026Updated 7 months ago
- Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning☆19Sep 25, 2025Updated last year