Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"
☆38Jun 3, 2026Updated 2 months ago
Alternatives and similar repositories for vstat
Users that are interested in vstat are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Nov 25, 2025Updated 8 months ago
- ReMoDetect: Reward Models Recognize Aligned LLM's Generations (NeurIPS 2024)☆17Nov 15, 2024Updated last year
- Coarse-to-fine Q-Network☆59Aug 6, 2024Updated last year
- Cambrian-S: Towards Spatial Supersensing in Video☆564Apr 3, 2026Updated 4 months ago
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code release for "PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop" (ICML 2025)☆59May 8, 2025Updated last year
- Subtask-Aware Visual Reward Learning from Segmented Demonstrations (ICLR 2025 accepted)☆19Apr 11, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆18Jun 2, 2026Updated 2 months ago
- PyTorch code accompanying the paper "Imitating Graph-Based Planning with Goal-Conditioned Policies" (ICLR 2023).☆21Mar 4, 2023Updated 3 years ago
- [ICML 2026] HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling☆29May 2, 2026Updated 3 months ago
- Meta-Learning with Self-Improving Momentum Target (NeurIPS 2022)☆23Oct 12, 2022Updated 3 years ago
- A comprehensive JAX/NNX library for diffusion and flow matching generative algorithms, featuring DiT (Diffusion Transformer) and its vari…☆152Oct 16, 2025Updated 9 months ago
- Official implementation of Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning (NeurIPS 2024).☆34Mar 4, 2025Updated last year
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆31Apr 2, 2026Updated 4 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆16Jun 19, 2026Updated last month
- Learning Large-scale Neural Fields via Context Pruned Meta-Learning (NeurIPS 2023)☆28Sep 24, 2023Updated 2 years ago
- ☆16Sep 11, 2025Updated 10 months ago
- Guide Your Agent with Adaptive Multimodal Rewards (NeurIPS 2023 Accepted)☆33Sep 25, 2023Updated 2 years ago
- ☆22Jul 26, 2026Updated last week
- Visual Representation Learning with Stochastic Frame Prediction (ICML 2024)☆28Nov 27, 2024Updated last year
- ☆38Feb 6, 2025Updated last year
- AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.☆297May 5, 2026Updated 2 months ago
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs☆39Sep 22, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆29Apr 30, 2024Updated 2 years ago
- [CVPR 2026] Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos☆29Mar 25, 2026Updated 4 months ago
- Contrastive self-supervised learning using Rényi divergence☆14Oct 21, 2022Updated 3 years ago
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence☆61Mar 11, 2026Updated 4 months ago
- [ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation☆16Jun 4, 2026Updated last month
- ☆19Nov 29, 2024Updated last year
- LogiCity@NeurIPS'24, D&B track. A multi-agent inductive learning environment for "abstractions".☆27Jun 10, 2025Updated last year
- Code for the paper "What Makes Better Augmentation Strategies? Augment Difficult but Not too Different" (ICLR 22)☆12Aug 28, 2023Updated 2 years ago
- Code for "Agentic Very Long Video Understanding" (EGAgent) [ACL 2026 Main]☆51Jul 1, 2026Updated last month
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆391Jun 20, 2026Updated last month
- ☆22Jun 10, 2024Updated 2 years ago
- Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders☆255Feb 13, 2026Updated 5 months ago
- A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.☆341Jul 27, 2026Updated last week
- ☆31Nov 18, 2025Updated 8 months ago
- Official implementation of Decoupled MeanFlow☆44Oct 28, 2025Updated 9 months ago
- PaperBot: Learning to Design Real-World Tools Using Paper☆13Mar 15, 2024Updated 2 years ago