[ICLR 2025] Video Action Differencing
β53Jul 3, 2025Updated last year
Alternatives and similar repositories for viddiff
Users that are interested in viddiff are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Repo for our work "Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence"β20Jun 2, 2025Updated last year
- [CVPR 2025] MicroVQA eval and π€RefineBot code for "MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research"β¦β36Nov 25, 2025Updated 9 months ago
- β41Sep 9, 2025Updated 11 months ago
- [EACL 2026] PaperSearchQA. Data generation pipeline for QA over scientific papers, suitable for RL training search agentsβ36Feb 4, 2026Updated 7 months ago
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Modelsβ25Mar 21, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2025] BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literatureβ108Mar 22, 2025Updated last year
- β57Jun 8, 2026Updated 2 months ago
- [ICLR 2025] Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervisionβ72Jul 10, 2024Updated 2 years ago
- Official implementation of "Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data" (ICLR 2024)β37Oct 16, 2024Updated last year
- [CVPR 2025] Custom Open CLIP repo to train biomedical CLIP modelsβ38Mar 23, 2025Updated last year
- Official implementation of "Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation" (CVPR 202β¦β40May 26, 2025Updated last year
- Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignβ¦β19Apr 5, 2024Updated 2 years ago
- [ECCV 2024] Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Modelsβ111Dec 3, 2024Updated last year
- Official Code Release for "Diagnosing and Rectifying Vision Models using Language" (ICLR 2023)β34Jun 8, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [CVPR 2023] Official PyTorch implementation of the paper "GAP: Post-Processing Temporal Action Detection"β18Aug 31, 2023Updated 3 years ago
- Official implementation of "Describing Differences in Image Sets with Natural Language" (CVPR 2024 Oral)β134Nov 5, 2025Updated 10 months ago
- Automated Qualitative Analysis of LLMs (ICLR 2025)β53Jul 6, 2025Updated last year
- [ICCV, 2023] Multiple humans in 3D captured by dynamic and static cameras in 4K.β52Nov 24, 2023Updated 2 years ago
- Data release for Step Differences in Instructional Video (CVPR24)β15Jun 19, 2024Updated 2 years ago
- Official code repository for "Video-Mined Task Graphs for Keystep Recognition in Instructional Videos" arXiv, 2023β15Apr 1, 2024Updated 2 years ago
- [NeurIPS 2023] Official Pytorch code for LOVM: Language-Only Vision Model Selectionβ21Feb 3, 2024Updated 2 years ago
- The official PyTorch implementation of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) '24 paper PREGO: online mistake detectβ¦β35Jun 9, 2025Updated last year
- Code for paper: Unified Text-to-Image Generation and Retrievalβ15Jul 19, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official implementation of "ConViS-Bench: Estimating Video Similarity Through Semantic Concepts", NeurIPS 2025β27Nov 28, 2025Updated 9 months ago
- [CVPR 2023] LOGO: A Long-Form Video Dataset for Group Action Quality Assessmentβ48Apr 9, 2024Updated 2 years ago
- (ACL 2025) MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scaleβ50Jun 4, 2025Updated last year
- COLA: Evaluate how well your vision-language model can Compose Objects Localized with Attributes!β25May 14, 2026Updated 3 months ago
- [CVPR 2025] HumanMM: Global Human Motion Recovery from Multi-shot Videosβ122Mar 20, 2025Updated last year
- Automatically Analyze your Model Tracesβ47Mar 16, 2026Updated 5 months ago
- β62Apr 28, 2025Updated last year
- [CVPR 2025] GPS as a Control Signal for Image Generationβ25Mar 18, 2025Updated last year
- Geometry-aware Novel View Synthesis with Pre-trained 2D Priorβ39Jun 3, 2023Updated 3 years ago
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- β21Jun 30, 2025Updated last year
- β28Dec 22, 2024Updated last year
- β14May 25, 2021Updated 5 years ago
- Official repository for the ACL 2025 Findings paper "Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Mβ¦β26May 12, 2026Updated 3 months ago
- Implementation of the Benchmark Approaches for Medical Instructional Video Classification (MedVidCL) and Medical Video Question Answeringβ¦β31Jan 31, 2023Updated 3 years ago
- β47Dec 10, 2021Updated 4 years ago
- vLLM client with minimal dependenciesβ15Feb 28, 2024Updated 2 years ago