Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"
☆48Aug 12, 2026Updated last month
Alternatives and similar repositories for vstat
Users that are interested in vstat are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Aug 9, 2026Updated last month
- Cambrian-P: Pose-Grounded Video Understanding☆117Jul 28, 2026Updated 2 months ago
- Jaehyung Kim et al's ACL 2023 paper on "infoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-informat…☆16Jun 28, 2023Updated 3 years ago
- Coarse-to-fine Q-Network☆59Aug 6, 2024Updated 2 years ago
- Cambrian-S: Towards Spatial Supersensing in Video☆570Apr 3, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- [ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning☆90Jul 8, 2026Updated 2 months ago
- Code release for "PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop" (ICML 2025)☆61May 8, 2025Updated last year
- Subtask-Aware Visual Reward Learning from Segmented Demonstrations (ICLR 2025 accepted)☆20Apr 11, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆20Jun 2, 2026Updated 4 months ago
- PyTorch code accompanying the paper "Imitating Graph-Based Planning with Goal-Conditioned Policies" (ICLR 2023).☆22Mar 4, 2023Updated 3 years ago
- The first multiplayer video world model in Minecraft☆230Mar 3, 2026Updated 6 months ago
- [ICML 2026] HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling☆31May 2, 2026Updated 5 months ago
- Meta-Learning with Self-Improving Momentum Target (NeurIPS 2022)☆23Oct 12, 2022Updated 3 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- A comprehensive JAX/NNX library for diffusion and flow matching generative algorithms, featuring DiT (Diffusion Transformer) and its vari…☆155Oct 16, 2025Updated 11 months ago
- ☆31Feb 10, 2025Updated last year
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆33Apr 2, 2026Updated 6 months ago
- ☆17Jun 19, 2026Updated 3 months ago
- ☆62Apr 16, 2023Updated 3 years ago
- Learning Large-scale Neural Fields via Context Pruned Meta-Learning (NeurIPS 2023)☆29Sep 24, 2023Updated 3 years ago
- ☆16Sep 11, 2025Updated last year
- Openpi for RoboCasa Benchmark☆20May 12, 2026Updated 4 months ago
- ☆38Feb 6, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Visual Representation Learning with Stochastic Frame Prediction (ICML 2024)☆28Nov 27, 2024Updated last year
- ☆15Aug 11, 2025Updated last year
- [CVPR 2025] A Hierarchical Movie Level Dataset for Long Video Generation☆103Mar 16, 2025Updated last year
- Jump to better conclusions: SCAN both left and right☆11Jan 24, 2019Updated 7 years ago
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs☆39Sep 22, 2025Updated last year
- [NeurIPS 2026] Official Implementation of "Visual-ERM: Reward Modeling for Visual Equivalence"☆67Mar 23, 2026Updated 6 months ago
- [CVPR 2026] Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos☆30Mar 25, 2026Updated 6 months ago
- Open Source Graph Neural Net Based Pipeline for Image Matching☆21Jun 6, 2025Updated last year
- Contrastive self-supervised learning using Rényi divergence☆14Oct 21, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [EMNLP 2025 Oral] IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents☆25Sep 16, 2025Updated last year
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence☆62Mar 11, 2026Updated 6 months ago
- [ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation☆16Jun 4, 2026Updated 3 months ago
- ☆19Nov 29, 2024Updated last year
- LogiCity@NeurIPS'24, D&B track. A multi-agent inductive learning environment for "abstractions".☆27Jun 10, 2025Updated last year
- Code for the paper "What Makes Better Augmentation Strategies? Augment Difficult but Not too Different" (ICLR 22)☆12Aug 28, 2023Updated 3 years ago
- Code for "Agentic Very Long Video Understanding" (EGAgent) [ACL 2026 Main]☆63Sep 1, 2026Updated last month