JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
☆1,934Sep 15, 2026Updated last week
Alternatives and similar repositories for JoyAI-VL-Interaction
Users that are interested in JoyAI-VL-Interaction are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds☆2,043Updated this week
- ☆157Apr 27, 2026Updated 4 months ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆401Jun 20, 2026Updated 3 months ago
- 🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖☆473Aug 5, 2026Updated last month
- [ICLR 2026] An official implementation of "STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence"☆45Apr 19, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Fully Open Framework for Democratized Multimodal Training☆1,212Updated this week
- 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for str…☆210Updated this week
- StreamingVLM: Real-Time Understanding for Infinite Video Streams☆1,085Oct 15, 2025Updated 11 months ago
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning☆37Aug 28, 2025Updated last year
- [ACM MM 2025] TimeChat-online: 80% Visual Tokens are Naturally Redundant in Streaming Videos☆132Jun 29, 2026Updated 2 months ago
- ☆127Jun 5, 2026Updated 3 months ago
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆183Aug 14, 2026Updated last month
- Implementation of Em_Garde: a proposal-retrieval framework for streaming video understanding☆33Jun 24, 2026Updated 3 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆20Apr 2, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated 2 months ago
- Streaming Video Instruction Tuning☆91Feb 25, 2026Updated 7 months ago
- [CVPR 2026] An official implementation of "Think Visually, Reason Textually: Vision-Language Synergy in ARC"☆47Nov 26, 2025Updated 9 months ago
- [ICML 2025 Oral] An official implementation of VideoRoPE & VideoRoPE++☆224Apr 15, 2026Updated 5 months ago
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆114Mar 14, 2025Updated last year
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,509Mar 9, 2026Updated 6 months ago
- [CVPR 2025] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?☆164Jul 24, 2025Updated last year
- A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Lan…☆23May 25, 2026Updated 4 months ago
- ☆1,453Feb 12, 2026Updated 7 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 7 months ago
- Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"☆426Jan 29, 2026Updated 7 months ago
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)☆479Oct 29, 2025Updated 10 months ago
- [NeurIPS 2025] Official implementation of HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance☆90Sep 18, 2025Updated last year
- A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.☆429Updated this week
- ☆122Jul 26, 2026Updated last month
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆169Jun 10, 2026Updated 3 months ago
- Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, im…☆4,029Apr 23, 2026Updated 5 months ago
- [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactiv…☆976Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence☆962Sep 17, 2026Updated last week
- [ECCV 2026] Offline implementation of UniREditBench: A Unified Reasoning-based Image Editing Benchmark.☆59Aug 14, 2026Updated last month
- JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image ed…☆2,157Aug 5, 2026Updated last month
- Multimodal RL training framework for diffusion & omni models☆1,106Updated this week
- [NeurIPS2025] The official PyTorch implementation of the "Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video".☆35Dec 25, 2025Updated 9 months ago
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning☆58Jun 2, 2026Updated 3 months ago
- Official implementation of BLIP3o-Series☆1,667Nov 29, 2025Updated 9 months ago