☆34Jun 2, 2023Updated 3 years ago
Alternatives and similar repositories for TranS4mer
Users that are interested in TranS4mer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆143Jan 3, 2024Updated 2 years ago
- This is an official PyTorch Implementation of Neighbor Relations Matter in Video Scene Detection.☆30Mar 19, 2025Updated last year
- Official PyTorch implementation of "Ground-A-Score: Scaling Up the Score Distillation for Multi-Attribute Editing"☆11Apr 4, 2024Updated 2 years ago
- Code for CVPR 2022 paper "Scene Consistency Representation Learning for Video Scene Segmentation"☆112Feb 14, 2023Updated 3 years ago
- Codebase for CVPR2020 A Local-to-Global Approach to Multi-modal Movie Scene Segmentation☆239May 20, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- [T-PAMI 2023] Temporal Perceiver: A General Architecture for Arbitrary Boundary Detection☆39Aug 29, 2023Updated 3 years ago
- PyTorch implementation of "PatchGame: Learning to Signal Mid-level Patches in Referential Games" to appear in NeurIPS 2021☆24Jun 4, 2021Updated 5 years ago
- Constraint Satisfaction Visual Grounding☆16Aug 10, 2025Updated last year
- [AAAI-24] VVS : Video-to-Video Retrieval With Irrelevant Frame Suppression☆21May 14, 2024Updated 2 years ago
- Code for OrthDNNs: Orthogonal Deep Neural Networks☆14Jan 9, 2020Updated 6 years ago
- ☆24Sep 24, 2023Updated 2 years ago
- Multi-modal transformer approach for natural language query based joint video summarization and highlight detection☆17May 23, 2024Updated 2 years ago
- TransNet V2: Shot Boundary Detection Neural Network☆1,034Dec 4, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for CVPR2023 paper "Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies"☆18Mar 21, 2023Updated 3 years ago
- A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot-level captions.☆180Jan 30, 2025Updated last year
- Character-aware audio-only subtitling☆31Jun 15, 2025Updated last year
- [CVPR21] Visual Semantic Role Labeling for Video Understanding (https://arxiv.org/abs/2104.00990)☆61Aug 17, 2021Updated 5 years ago
- Official code for "Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models" (TPDM)☆67Jun 26, 2024Updated 2 years ago
- ☆18Aug 19, 2024Updated 2 years ago
- [CVPR2024] The official implementation of AdaTAD: End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames☆43Jul 9, 2024Updated 2 years ago
- Repo from the "Learning with limited labeled data" seminar @ Uni of Tuebingen. A collection of notes, notebooks and slideshows to underst…☆17Apr 13, 2023Updated 3 years ago
- windows setup script☆11Jan 22, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [IEEE T-IP 2022] TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning☆25Dec 19, 2023Updated 2 years ago
- Code for the paper: Graph Jigsaw Learning for Cartoon Face Recognition☆10Jul 1, 2022Updated 4 years ago
- Provably (and non-vacuously) bounding test error of deep neural networks under distribution shift with unlabeled test data.☆10Feb 27, 2024Updated 2 years ago
- ☆15Feb 13, 2025Updated last year
- PyTorch implementation of "PatchVAE: Learning Local Latent Codes for Recognition" to appear in CVPR 2020☆14Apr 9, 2020Updated 6 years ago
- Code for DVD A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue☆14Oct 12, 2021Updated 4 years ago
- 为视障人群生成电影,输入是电影剧本和mkv格式电影,输出为带有解说的电影☆12Jul 28, 2019Updated 7 years ago
- A guide to structured generation using constrained decoding☆18Jun 9, 2024Updated 2 years ago
- [ICIP2023] Code for the paper 'Action Anticipation with Goal Consistency'☆12Apr 5, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- FBNet code for FGOC in aerial images☆15Jun 9, 2022Updated 4 years ago
- ☆85Mar 10, 2025Updated last year
- Quick Long Video Understanding [TMLR2025]☆78Oct 27, 2025Updated 10 months ago
- [NeurIPS 2023 D&B] VidChapters-7M: Video Chapters at Scale☆214Nov 13, 2023Updated 2 years ago
- A series of Jupyter notebooks that walk you through the fundamentals of Machine Learning and Deep Learning in python using Scikit-Learn a…☆11Jun 8, 2019Updated 7 years ago
- 【CVPR'24】OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition☆39Apr 27, 2024Updated 2 years ago
- Generate interleaved text and image content in a structured format you can directly pass to downstream APIs.☆29Oct 18, 2024Updated last year