JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
☆1,849Aug 29, 2026Updated last week
Alternatives and similar repositories for JoyAI-VL-Interaction
Users that are interested in JoyAI-VL-Interaction are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds☆2,002Updated this week
- ☆155Apr 27, 2026Updated 4 months ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆397Jun 20, 2026Updated 2 months ago
- 🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖☆458Aug 5, 2026Updated last month
- [ICLR 2026] An official implementation of "STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence"☆44Apr 19, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Fully Open Framework for Democratized Multimodal Training☆1,197Updated this week
- 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for str…☆198Aug 17, 2026Updated 2 weeks ago
- StreamingVLM: Real-Time Understanding for Infinite Video Streams☆1,077Oct 15, 2025Updated 10 months ago
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning☆37Aug 28, 2025Updated last year
- [ACM MM 2025] TimeChat-online: 80% Visual Tokens are Naturally Redundant in Streaming Videos☆133Jun 29, 2026Updated 2 months ago
- ☆125Jun 5, 2026Updated 3 months ago
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆182Aug 14, 2026Updated 3 weeks ago
- Implementation of Em_Garde: a proposal-retrieval framework for streaming video understanding☆30Jun 24, 2026Updated 2 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆20Apr 2, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated last month
- Streaming Video Instruction Tuning☆87Feb 25, 2026Updated 6 months ago
- [CVPR 2026] An official implementation of "Think Visually, Reason Textually: Vision-Language Synergy in ARC"☆47Nov 26, 2025Updated 9 months ago
- [ICML 2025 Oral] An official implementation of VideoRoPE & VideoRoPE++☆224Apr 15, 2026Updated 4 months ago
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆112Mar 14, 2025Updated last year
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,502Mar 9, 2026Updated 5 months ago
- [CVPR 2025] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?☆161Jul 24, 2025Updated last year
- A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Lan…☆23May 25, 2026Updated 3 months ago
- ☆1,443Feb 12, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 7 months ago
- Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"☆425Jan 29, 2026Updated 7 months ago
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)☆477Oct 29, 2025Updated 10 months ago
- [NeurIPS 2025] Official implementation of HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance☆90Sep 18, 2025Updated 11 months ago
- A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.☆414Aug 25, 2026Updated last week
- ☆110Jul 26, 2026Updated last month
- Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, im…☆4,000Apr 23, 2026Updated 4 months ago
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆168Jun 10, 2026Updated 2 months ago
- [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactiv…☆947Aug 28, 2026Updated last week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence☆944Aug 5, 2026Updated last month
- [ECCV 2026] Offline implementation of UniREditBench: A Unified Reasoning-based Image Editing Benchmark.☆58Aug 14, 2026Updated 3 weeks ago
- JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image ed…☆2,150Aug 5, 2026Updated last month
- [NeurIPS2025] The official PyTorch implementation of the "Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video".☆35Dec 25, 2025Updated 8 months ago
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning☆57Jun 2, 2026Updated 3 months ago
- Official implementation of BLIP3o-Series☆1,667Nov 29, 2025Updated 9 months ago
- DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action☆136May 20, 2026Updated 3 months ago