A benchmark for evaluating future-event forecasting from audio and video context in multimodal language models
☆29Jan 22, 2026Updated 7 months ago
Alternatives and similar repositories for FutureOmni
Users that are interested in FutureOmni are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A multi-user OpenClaw deployment with isolated workspaces and configurable model backends☆25Mar 30, 2026Updated 5 months ago
- ☆27Jan 29, 2026Updated 7 months ago
- A python tool help to interact with chatgpt.☆10Dec 11, 2022Updated 3 years ago
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs☆51Jul 12, 2026Updated 2 months ago
- FamilyTool benchmark☆14Sep 10, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆146Aug 31, 2026Updated 2 weeks ago
- We introduce 'Thinking with Video', a new paradigm leveraging video generation for multimodal reasoning. Our VideoThinkBench shows that S…☆320Aug 23, 2026Updated 3 weeks ago
- Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication☆21Mar 21, 2024Updated 2 years ago
- Code and data release for the paper "Seeing the Arrow of Time in Large Multimodal Models"☆17Oct 2, 2025Updated 11 months ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models☆108Apr 20, 2026Updated 5 months ago
- A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical …☆61Sep 1, 2025Updated last year
- Source code for our paper ''Finding What Matters: Anchoring Context Knowledge with Evolving Indices for Iterative Retrieval''☆28Jun 2, 2026Updated 3 months ago
- [EMNLP 2026 Main] Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding☆28Aug 21, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- We propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervis…☆432Updated this week
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 7 months ago
- A proactive robot manipulation model for multimodal physical environments☆120Mar 28, 2026Updated 5 months ago
- The official implementation of the paper "MotifRetro: Exploring the Combinability-Consistency Trade-offs in retrosynthesis via Dynamic Mo…☆11Jun 25, 2023Updated 3 years ago
- A post-training framework for diffusion language models with supervised fine-tuning and reinforcement learning☆165Mar 30, 2026Updated 5 months ago
- Source code for the NAACL 2021 paper: "Distantly Supervised Relation Extraction with Sentence Reconstruction and Knowledge Base Priors"☆12Jul 15, 2021Updated 5 years ago
- ☆18Aug 1, 2024Updated 2 years ago
- [ACL2025 main] Official implementation of "LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjo…☆20Mar 16, 2026Updated 6 months ago
- ☆34Mar 17, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ACL 2026 Findings] "Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning"☆63May 26, 2026Updated 3 months ago
- [CVPR 2026] Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding☆33Apr 12, 2026Updated 5 months ago
- Public codebase for ECONET: EMNLP'21☆12Mar 11, 2022Updated 4 years ago
- Code for paper: “What Data Benefits My Classifier?” Enhancing Model Performance and Interpretability through Influence-Based Data Selecti…☆23May 17, 2024Updated 2 years ago
- ☆20Jan 29, 2026Updated 7 months ago
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23Sep 4, 2026Updated 2 weeks ago
- ☆29Apr 8, 2025Updated last year
- A foundation model that generates synchronized video and audio in a single model☆1,117Updated this week
- ☆14Feb 26, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆20Apr 2, 2025Updated last year
- Some simple tutorials about python☆12Oct 11, 2020Updated 5 years ago
- On Path to Multimodal Generalist: General-Level and General-Bench☆22Jul 11, 2025Updated last year
- Diffusion-based generative drug-like molecular editing with chemical natural language☆18Dec 22, 2024Updated last year
- ☆32Feb 27, 2025Updated last year
- \infty-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation☆22Feb 14, 2025Updated last year
- TTRV: Test-Time Reinforcement Learning for Vision–Language Models (CVPR 2026)☆47Mar 8, 2026Updated 6 months ago