Multi-model video-to-text by combining embeddings from Flan-T5 + CLIP + Whisper + SceneGraph. The 'backbone LLM' is pre-trained from scratch on YouTube (YT-1B dataset).
☆54Apr 21, 2023Updated 3 years ago
Alternatives and similar repositories for video-pretrained-transformer
Users that are interested in video-pretrained-transformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ECCV2022] A PyTorch implementation of the paper "Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embo…☆13Mar 20, 2023Updated 3 years ago
- [NeurIPS 2023 D&B] VidChapters-7M: Video Chapters at Scale☆213Nov 13, 2023Updated 2 years ago
- pytorch implementation of Semantics-AssistedVideoCaptioning☆11Feb 16, 2023Updated 3 years ago
- Implementation of the model: "(MC-ViT)" from the paper: "Memory Consolidation Enables Long-Context Video Understanding"☆27Updated this week
- ☆17Aug 6, 2021Updated 5 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Implementation of VisionLLaMA from the paper: "VisionLLaMA: A Unified LLaMA Interface for Vision Tasks" in PyTorch and Zeta☆15Nov 11, 2024Updated last year
- Code for "Visual Spatial Description: Controlled Spatial-Oriented Image-to-Text Generation"☆25Mar 9, 2024Updated 2 years ago
- Implementation of DropCov as described in DropCov: A Simple yet Effective Method for Improving Deep Architectures☆10Oct 15, 2022Updated 3 years ago
- Implementation of a Hierarchical Mamba as described in the paper: "Hierarchical State Space Models for Continuous Sequence-to-Sequence Mo…☆16Nov 11, 2024Updated last year
- ☆58Apr 24, 2024Updated 2 years ago
- Super Resolution example☆10Oct 5, 2015Updated 10 years ago
- Additional utility code for the Spring dataset and benchmark☆13Jul 4, 2023Updated 3 years ago
- Source code of SACD(Super-resolution with Auto-Correlation two-step Deconvolution)☆13Dec 12, 2022Updated 3 years ago
- 32 times longer context window than vanilla Transformers and up to 4 times longer than memory efficient Transformers.☆51Jun 16, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ECCV 2024🔥] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"☆155Sep 10, 2024Updated last year
- GUI for cropping a large amount of images quickly.☆13May 24, 2018Updated 8 years ago
- CRCNet for image deblurring☆10Feb 6, 2019Updated 7 years ago
- Official implementation for paper Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos☆28Dec 8, 2023Updated 2 years ago
- Offline-first, decentralized graph database of collaborative Web apps☆15May 12, 2017Updated 9 years ago
- ☆18Aug 19, 2024Updated 2 years ago
- A much powerful probing method to tune your model with promising performance and linear probing training cost!☆15Jul 26, 2023Updated 3 years ago
- [CVPR 2024] Do you remember? Dense Video Captioning with Cross-Modal Memory Retrieval☆67Jun 19, 2024Updated 2 years ago
- Ultra Fast Multi-Modality Vector Database☆18Feb 21, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official code for the LoG2022 paper -- MSGNN: A Spectral Graph Neural Network Based on a Novel Magnetic Signed Laplacian.☆14Jul 30, 2026Updated last month
- Learning Interactions and Relationships between Movie Characters (CVPR'20)☆22Apr 12, 2023Updated 3 years ago
- Audio-visual diarization pipeline used for creating VoxConverse dataset☆22Jun 6, 2025Updated last year
- Multimodal and multilingual topic model with pretrained embeddings☆12Apr 11, 2023Updated 3 years ago
- Implementation of MambaFormer in Pytorch ++ Zeta from the paper: "Can Mamba Learn How to Learn? A Comparative Study on In-Context Learnin…☆23Updated this week
- Leveraging DSPy for AI-driven task understanding and solution generation, the Self-Discover Framework automates problem-solving through r…☆74Nov 4, 2025Updated 9 months ago
- Simulates agent path planning using A* and Q-Learning in a 2D grid☆12Apr 5, 2014Updated 12 years ago
- Демонстрация структуры ml проекта☆11Oct 12, 2022Updated 3 years ago
- [WACV 2026 Oral] LASER: Lip Landmark Assisted Speaker Detection for Robustness official implemntation☆30Feb 26, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Task-Focused Few-Shot Object Detection Benchmark☆14Jun 24, 2025Updated last year
- Generate interleaved text and image content in a structured format you can directly pass to downstream APIs.☆29Oct 18, 2024Updated last year
- End-to-End Dense Video Captioning with Parallel Decoding (ICCV 2021)☆230Jan 3, 2024Updated 2 years ago
- voxelnet on ASTYX radar dataset for automotive object detection☆13Aug 28, 2020Updated 6 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- Train punctuation and capitalization models for different languages☆26Apr 2, 2022Updated 4 years ago
- [ICCV2023 Oral] Implicit Temporal Modeling with Learnable Alignment for Video Recognition☆41Nov 29, 2023Updated 2 years ago