[ICLR 2025] Video Action Differencing
β54Jul 3, 2025Updated last year
Alternatives and similar repositories for viddiff
Users that are interested in viddiff are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Repo for our work "Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence"β21Updated this week
- [CVPR 2025] MicroVQA eval and π€RefineBot code for "MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research"β¦β36Nov 25, 2025Updated 9 months ago
- β15Jan 7, 2026Updated 8 months ago
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Modelsβ26Mar 21, 2026Updated 6 months ago
- [EACL 2026] PaperSearchQA. Data generation pipeline for QA over scientific papers, suitable for RL training search agentsβ37Feb 4, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A Vision-Language Benchmark for Microscopy Understandingβ31Mar 13, 2025Updated last year
- [ICLR 2025] Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervisionβ72Jul 10, 2024Updated 2 years ago
- Code implementation of the paper 'ExpertAF: Expert Actionable Feedback from Video'β20Sep 30, 2025Updated 11 months ago
- Official implementation of "Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data" (ICLR 2024)β37Oct 16, 2024Updated last year
- [CVPR 2025] Custom Open CLIP repo to train biomedical CLIP modelsβ38Mar 23, 2025Updated last year
- Official implementation of "Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation" (CVPR 202β¦β39May 26, 2025Updated last year
- Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignβ¦β19Apr 5, 2024Updated 2 years ago
- [ECCV 2024] Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Modelsβ112Dec 3, 2024Updated last year
- Official Code Release for "Diagnosing and Rectifying Vision Models using Language" (ICLR 2023)β34Jun 8, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [NeurIPS 2024] Official implementation of paper "Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers"β22Mar 10, 2025Updated last year
- Official implementation of "Describing Differences in Image Sets with Natural Language" (CVPR 2024 Oral)β135Nov 5, 2025Updated 10 months ago
- [COLING 2022]: CommunityLM: Probing Partisan Worldviews from Language Modelsβ14Jan 31, 2023Updated 3 years ago
- Automated Qualitative Analysis of LLMs (ICLR 2025)β54Jul 6, 2025Updated last year
- Data release for Step Differences in Instructional Video (CVPR24)β15Jun 19, 2024Updated 2 years ago
- Official code repository for "Video-Mined Task Graphs for Keystep Recognition in Instructional Videos" arXiv, 2023β15Apr 1, 2024Updated 2 years ago
- The QA datasets used for DrQA evaluation.β14Nov 30, 2018Updated 7 years ago
- [NeurIPS 2023] Official Pytorch code for LOVM: Language-Only Vision Model Selectionβ21Feb 3, 2024Updated 2 years ago
- The official PyTorch implementation of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) '24 paper PREGO: online mistake detectβ¦β35Jun 9, 2025Updated last year
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Does patch ordering affect context-limited vision transformers?β17Oct 10, 2025Updated 11 months ago
- Code for paper: Unified Text-to-Image Generation and Retrievalβ15Jul 19, 2026Updated 2 months ago
- Official implementation of "Why are Visually-Grounded Language Models Bad at Image Classification?" (NeurIPS 2024)β99Oct 19, 2024Updated last year
- Official implementation of "ConViS-Bench: Estimating Video Similarity Through Semantic Concepts", NeurIPS 2025β28Nov 28, 2025Updated 9 months ago
- An official implementation of Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformersβ13Mar 9, 2023Updated 3 years ago
- [CVPR 2023] LOGO: A Long-Form Video Dataset for Group Action Quality Assessmentβ48Apr 9, 2024Updated 2 years ago
- (ACL 2025) MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scaleβ50Jun 4, 2025Updated last year
- COLA: Evaluate how well your vision-language model can Compose Objects Localized with Attributes!β25May 14, 2026Updated 4 months ago
- CoMA: Compositional Human Motion Generation with Multi-modal Agentsβ17Sep 9, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2025] HumanMM: Global Human Motion Recovery from Multi-shot Videosβ123Mar 20, 2025Updated last year
- Automatically Analyze your Model Tracesβ47Mar 16, 2026Updated 6 months ago
- β62Apr 28, 2025Updated last year
- [CVPR 2025] GPS as a Control Signal for Image Generationβ25Mar 18, 2025Updated last year
- Geometry-aware Novel View Synthesis with Pre-trained 2D Priorβ39Jun 3, 2023Updated 3 years ago
- Repository for 3DV2022 paper "Domain Adaptive 3D Pose Augmentation for In-the-wild Human Mesh Recovery"β19Mar 22, 2023Updated 3 years ago
- β14May 25, 2021Updated 5 years ago