Code and data release for the paper "Learning Object State Changes in Videos: An Open-World Perspective" (CVPR 2024)
β37Sep 9, 2024Updated last year
Alternatives and similar repositories for VidOSC
Users that are interested in VidOSC are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official code repository for "Video-Mined Task Graphs for Keystep Recognition in Instructional Videos" arXiv, 2023β15Apr 1, 2024Updated 2 years ago
- π₯π₯π₯ Object State Description & Change Detectionβ10Apr 6, 2026Updated 4 months ago
- The official PyTorch implementation of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) '24 paper PREGO: online mistake detectβ¦β35Jun 9, 2025Updated last year
- Code release for the paper "Egocentric Video Task Translation" (CVPR 2023 Highlight)β34Jun 12, 2023Updated 3 years ago
- ECCV 2024 STMA & CVPR 2024 1st MOSE & 1st VOT Challenge & 1st LSVOS v6β12Oct 16, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official PyTorch code of GroundVQA (CVPR'24)β63Sep 13, 2024Updated last year
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task Videosβ38Sep 9, 2024Updated last year
- [CVPR 2024] KEPP: Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videosβ12Sep 24, 2024Updated last year
- Data release for Step Differences in Instructional Video (CVPR24)β15Jun 19, 2024Updated 2 years ago
- [ICCV 2023] How Much Temporal Long-Term Context is Needed for Action Segmentation?β50Jun 21, 2024Updated 2 years ago
- (NeurIPS 2023) Open-set visual object query search & localization in long-form videosβ26Feb 1, 2024Updated 2 years ago
- Code implementation for our ECCV, 2022 paper titled "My View is the Best View: Procedure Learning from Egocentric Videos"β35Feb 5, 2024Updated 2 years ago
- β31Mar 5, 2025Updated last year
- [CVPR 2024 Champions][ICLR 2025] Solutions for EgoVis Chanllenges in CVPR 2024β136May 11, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The champion solution for Ego4D Natural Language Queries Challenge in CVPR 2023β18Jan 23, 2024Updated 2 years ago
- β45Jan 13, 2026Updated 7 months ago
- For Ego4D VQ3D Taskβ22Jan 9, 2024Updated 2 years ago
- HT-Step is a large-scale article grounding dataset of temporal step annotations on how-to videosβ26Mar 20, 2024Updated 2 years ago
- Code release for the paper "Progress-Aware Video Frame Captioning" (CVPR 2025)β26Jul 16, 2025Updated last year
- β25Jul 10, 2023Updated 3 years ago
- Code for CVPR 2023 paper "Procedure-Aware Pretraining for Instructional Video Understanding"β50Jun 2, 2026Updated 3 months ago
- Code for ECCV2022 "Real-time Online Video Detection with Temporal Smoothing Transformers"β119Aug 23, 2025Updated last year
- Code for the paper "ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions" published at CVPR 2025β24Mar 16, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for the paper "GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos" published at CVPR 2024β54Mar 3, 2024Updated 2 years ago
- Human-centric environment representations from egocentric videoβ15Feb 5, 2026Updated 6 months ago
- Official implementation of GROOT, CoRL 2023β71Nov 4, 2023Updated 2 years ago
- β11Apr 16, 2023Updated 3 years ago
- Repo for paper: "Paxion: Patching Action Knowledge in Video-Language Foundation Models" Neurips 23 Spotlightβ38May 23, 2023Updated 3 years ago
- [TPAMI'2023]Knowledge-enriched Attention Network with Group-wise Semantic for Visual Storytellingβ11Jan 3, 2023Updated 3 years ago
- A PyTorch implementation of the paper "MMoT: Mixture-of-Modality-Tokens Transformer for Composed Multimodal Conditional Image Synthesis".β12Jan 16, 2023Updated 3 years ago
- [ICLR 2024] Towards Robust Multi-Modal Reasoning via Model Selectionβ14Mar 7, 2024Updated 2 years ago
- [ECCV 2024] VISAGE: Video Instance Segmentation with Appearance-Guided Enhancementβ39Jul 29, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- β31Aug 19, 2024Updated 2 years ago
- Official implementation of `Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning`, CVPR 2025β13Updated this week
- X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization, CVPR 2024β11Nov 7, 2024Updated last year
- Reproducing the paper "The continuous Bernoulli: fixing a pervasive error in variational autoencoders" for the Reproducibility Challenge β¦β12Jul 25, 2024Updated 2 years ago
- [CVPRW'23 Best Paper Award] Zero-shot Unsupervised Transfer Instance Segmentationβ24Aug 22, 2023Updated 3 years ago
- Get CLIP ViT text tokens about an image, visualize attention as a heatmap.β15Aug 8, 2023Updated 3 years ago
- β82Jan 5, 2024Updated 2 years ago