This repository provides core code for managing large volumes of video footage, enabling content understanding, automatic tagging, and vector database storage. It integrates multimodal models and LLMs for accurate descriptions and semantic search. A web interface allows visualization.
☆21Mar 25, 2025Updated last year
Alternatives and similar repositories for video-understanding
Users that are interested in video-understanding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆10Apr 20, 2023Updated 3 years ago
- ☆14Feb 26, 2024Updated 2 years ago
- 0xAA Wallet is a AA (Account Abstraction) wallet focused on developer experience, which helps developers build ERC4337 compatible Dapp.☆11Apr 1, 2023Updated 3 years ago
- code for paper Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering☆14Aug 13, 2024Updated 2 years ago
- ☆15Aug 12, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- for Mac OS X NSScrollView Pull to Refresh☆14Dec 23, 2015Updated 10 years ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model☆15Apr 7, 2026Updated 4 months ago
- ☆11Oct 28, 2022Updated 3 years ago
- Zicx's Notebook.☆10Nov 7, 2025Updated 9 months ago
- 主要是使用ffmpeg和opengl实现实现一些video effect.而最初的目的主要是用于调试gl_transiton项目中的一些opengl效果而开发的一个demo☆10Dec 20, 2019Updated 6 years ago
- Core Foundation Lite for Android☆15Jun 1, 2019Updated 7 years ago
- ☆16Jan 12, 2026Updated 7 months ago
- ☆22Apr 24, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Trying to get my wired Xbox One controller to work.☆16Jun 4, 2025Updated last year
- 红外和可见光融合☆10Apr 17, 2019Updated 7 years ago
- Code for the CVPR'23 paper: "STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition"☆21Dec 9, 2024Updated last year
- Find strongest response of convolutional layers on an image dataset. Automatically compute receptive field for any CNN layer.☆14Feb 19, 2021Updated 5 years ago
- PyTorch feature detection -> feature matching -> multi-image homography -> panorama stitching algorithm☆12Dec 28, 2022Updated 3 years ago
- An open source cpu-based image processing☆15Mar 4, 2019Updated 7 years ago
- ☆18Jan 16, 2026Updated 6 months ago
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated 2 years ago
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Multigranularity Contrastive cross-modal collaborative Generation (MCG) model for Video QA☆12Dec 13, 2023Updated 2 years ago
- TonyCrane's Python Lecture☆13Nov 4, 2022Updated 3 years ago
- ☆10Jul 23, 2023Updated 3 years ago
- Code for RACE.☆15Nov 12, 2025Updated 9 months ago
- Docker version API for MODNet-model Human Matting☆16Dec 3, 2023Updated 2 years ago
- Medical System as a Service - ZJU Software Engineering Course Project 2021☆14Jul 7, 2021Updated 5 years ago
- ☆24Oct 31, 2023Updated 2 years ago
- Java swing + socket + mysql 五子棋网络对战游戏☆15Jan 26, 2020Updated 6 years ago
- ☆19Jul 22, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Course notes & Xminds of College of Software Technology, Zhejiang University.☆12Apr 20, 2019Updated 7 years ago
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆18Jun 2, 2026Updated 2 months ago
- ☆20Nov 8, 2023Updated 2 years ago
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23May 18, 2026Updated 2 months ago
- ☆17Oct 23, 2023Updated 2 years ago
- A simple game engine made with opengl☆13Aug 4, 2022Updated 4 years ago
- On Path to Multimodal Generalist: General-Level and General-Bench☆21Jul 11, 2025Updated last year