这是一个基于Pytorch平台、Transformer框架实现的视频描述生成 (Video Captioning) 深度学习模型。 视频描述生成任务指的是:输入一个视频,输出一句描述整个视频内容的文字(前提是视频较短且可以用一句话来描述)。本repo主要目的是帮助视力障碍者欣赏网络视频、感知周围环境,促进“无障碍视频”的发展。
☆100Mar 12, 2022Updated 4 years ago
Alternatives and similar repositories for Video-Captioning-Transformer
Users that are interested in Video-Captioning-Transformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The first unofficial implementation of CLIP4Caption: CLIP for Video Caption (ACMMM 2021)☆16Jan 2, 2023Updated 3 years ago
- [ICLR 26] This repository contains the code and dataset for our paper: Language-guided Open-world Video Anomaly Detection under Weak Supe…☆30May 7, 2026Updated 2 months ago
- The code of IJCAI22 paper "GL-RG: Global-Local Representation Granularity for Video Captioning".☆18May 10, 2023Updated 3 years ago
- The PyTorch code of the AAAI2021 paper "Non-Autoregressive Coarse-to-Fine Video Captioning".☆57Oct 22, 2023Updated 2 years ago
- This repository contains the code for a video captioning system inspired by Sequence to Sequence -- Video to Text. This system takes as i…☆170Oct 12, 2019Updated 6 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Source code for "Bi-modal Transformer for Dense Video Captioning" (BMVC 2020)☆231Apr 8, 2023Updated 3 years ago
- PyTorch implementation of Multi-modal Dense Video Captioning (CVPR 2020 Workshops)☆144Apr 8, 2023Updated 3 years ago
- Video Grounding and Captioning☆331Oct 12, 2021Updated 4 years ago
- Research code for CVPR 2022 paper "SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning"☆251May 26, 2022Updated 4 years ago
- 利用文字信息生成文字动画视频☆17Apr 14, 2022Updated 4 years ago
- Pytorch implementation of audio-visual fusion video captioning model☆27Jul 26, 2018Updated 7 years ago
- 自动生成短视频,文章自动成片,多模态混剪,数字人,声音克隆☆13Jun 25, 2024Updated 2 years ago
- Video captioning baseline models on Video2Commonsense Dataset.☆56Apr 15, 2021Updated 5 years ago
- Unofficial Pytorch Implementation of UNet3Plus: A Full-Scale Connected UNet for Medical Image Segmentation☆15Jul 22, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A Pytorch implementation of "Reconstruction Network for Video Captioning", CVPR 2018☆53Apr 6, 2020Updated 6 years ago
- ☆37Jan 20, 2023Updated 3 years ago
- pytorch implementation of Semantics-AssistedVideoCaptioning☆11Feb 16, 2023Updated 3 years ago
- This is the official implementation of ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos☆47Nov 5, 2025Updated 8 months ago
- ☆10Apr 20, 2018Updated 8 years ago
- A short course of visual modeling☆30Oct 9, 2025Updated 9 months ago
- (ACL'2023) MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning☆36Aug 8, 2024Updated last year
- [ICPR 2024] Exemplar-free continual deepfake detector that leverages CLIP and domain-specific multi-modal prompts☆15Aug 1, 2024Updated last year
- ☆17Aug 6, 2021Updated 4 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Simple High performance Infrastructure for Neural network Experiments☆14Sep 25, 2023Updated 2 years ago
- Hide text informations using invisible text characters☆12Apr 4, 2022Updated 4 years ago
- The official implementation of the paper **LVChat: Facilitating Long Video Comprehension**☆14Apr 15, 2024Updated 2 years ago
- A pytorch implementation of our paper Image Captioning with Inherent Sentiment (ICME 2021 Oral).☆11Jul 18, 2022Updated 4 years ago
- Patch-Diffusion Code (AAAI2022)☆13Mar 3, 2022Updated 4 years ago
- converting the pretrained tensorflow SoundNet model to pytorch☆14Jun 15, 2022Updated 4 years ago
- An official implementation for " UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation"☆366Jul 25, 2024Updated 2 years ago
- Flutter下拉选择菜单☆13Jan 30, 2023Updated 3 years ago
- Video summarization using Vision Transformers☆14Feb 5, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- kidney function classification and prediction through ultrasound-based kidney imaging: from deep learning to mass screening of chronic ki…☆12May 20, 2019Updated 7 years ago
- JSON Processing of RICO Dataset☆15Sep 16, 2022Updated 3 years ago
- ☆12Mar 3, 2025Updated last year
- Video Steganography / Watermarking Demo Tool☆10Jan 16, 2022Updated 4 years ago
- A VBPR implement by tensorflow☆17Aug 12, 2019Updated 6 years ago
- [Nature Machine Intelligence 2025] Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception☆154May 27, 2026Updated last month
- ☆12Oct 28, 2023Updated 2 years ago