Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"
☆40Aug 12, 2026Updated last week
Alternatives and similar repositories for vstat
Users that are interested in vstat are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Aug 9, 2026Updated 2 weeks ago
- ReMoDetect: Reward Models Recognize Aligned LLM's Generations (NeurIPS 2024)☆17Nov 15, 2024Updated last year
- Jaehyung Kim et al's ACL 2023 paper on "infoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-informat…☆16Jun 28, 2023Updated 3 years ago
- Coarse-to-fine Q-Network☆59Aug 6, 2024Updated 2 years ago
- Cambrian-S: Towards Spatial Supersensing in Video☆567Apr 3, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- [ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning☆89Jul 8, 2026Updated last month
- [NeurIPS'25 Spotlight] MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation☆20Feb 23, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆19Jun 2, 2026Updated 2 months ago
- PyTorch code accompanying the paper "Imitating Graph-Based Planning with Goal-Conditioned Policies" (ICLR 2023).☆21Mar 4, 2023Updated 3 years ago
- The first multiplayer video world model in Minecraft☆223Mar 3, 2026Updated 5 months ago
- [ICML 2026] HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling☆30May 2, 2026Updated 3 months ago
- Meta-Learning with Self-Improving Momentum Target (NeurIPS 2022)☆23Oct 12, 2022Updated 3 years ago
- A comprehensive JAX/NNX library for diffusion and flow matching generative algorithms, featuring DiT (Diffusion Transformer) and its vari…☆154Oct 16, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆31Apr 2, 2026Updated 4 months ago
- ☆16Jun 19, 2026Updated 2 months ago
- ☆62Apr 16, 2023Updated 3 years ago
- Learning Large-scale Neural Fields via Context Pruned Meta-Learning (NeurIPS 2023)☆28Sep 24, 2023Updated 2 years ago
- ☆16Sep 11, 2025Updated 11 months ago
- Guide Your Agent with Adaptive Multimodal Rewards (NeurIPS 2023 Accepted)☆33Sep 25, 2023Updated 2 years ago
- Openpi for RoboCasa Benchmark☆18May 12, 2026Updated 3 months ago
- ☆23Aug 14, 2026Updated last week
- Visual Representation Learning with Stochastic Frame Prediction (ICML 2024)☆28Nov 27, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆38Feb 6, 2025Updated last year
- ☆17Apr 14, 2026Updated 4 months ago
- ☆15Aug 11, 2025Updated last year
- AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.☆302May 5, 2026Updated 3 months ago
- [CVPR 2025] A Hierarchical Movie Level Dataset for Long Video Generation☆102Mar 16, 2025Updated last year
- Jump to better conclusions: SCAN both left and right☆11Jan 24, 2019Updated 7 years ago
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs☆39Sep 22, 2025Updated 11 months ago
- [CVPR 2026] Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos☆29Mar 25, 2026Updated 4 months ago
- Contrastive self-supervised learning using Rényi divergence☆14Oct 21, 2022Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence☆62Mar 11, 2026Updated 5 months ago
- [ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation☆16Jun 4, 2026Updated 2 months ago
- ☆19Nov 29, 2024Updated last year
- Code for the paper "What Makes Better Augmentation Strategies? Augment Difficult but Not too Different" (ICLR 22)☆12Aug 28, 2023Updated 2 years ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆396Jun 20, 2026Updated 2 months ago
- ☆22Jun 10, 2024Updated 2 years ago
- Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders☆258Feb 13, 2026Updated 6 months ago