This repository provides core code for managing large volumes of video footage, enabling content understanding, automatic tagging, and vector database storage. It integrates multimodal models and LLMs for accurate descriptions and semantic search. A web interface allows visualization.
☆21Mar 25, 2025Updated last year
Alternatives and similar repositories for video-understanding
Users that are interested in video-understanding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Feb 26, 2024Updated 2 years ago
- code for paper Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering☆14Aug 13, 2024Updated 2 years ago
- for Mac OS X NSScrollView Pull to Refresh☆14Dec 23, 2015Updated 10 years ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- An unnecessarily tiny and minimal implementation of GPT-2 in NumPy.☆11Feb 12, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆17Oct 25, 2024Updated last year
- Core Foundation Lite for Android☆15Jun 1, 2019Updated 7 years ago
- ☆15Jan 12, 2026Updated 7 months ago
- 红外和可见光融合☆10Apr 17, 2019Updated 7 years ago
- ☆13Oct 19, 2021Updated 4 years ago
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated 2 years ago
- Multigranularity Contrastive cross-modal collaborative Generation (MCG) model for Video QA☆12Dec 13, 2023Updated 2 years ago
- Code for RACE.☆15Nov 12, 2025Updated 9 months ago
- Docker version API for MODNet-model Human Matting☆16Dec 3, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Video action classification benchmark for common CNN architectures, implemented in PyTorch☆12Jan 31, 2022Updated 4 years ago
- code of cvpr26 paper Symphony☆17Apr 7, 2026Updated 4 months ago
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆19Jun 2, 2026Updated 2 months ago
- ☆15Jan 11, 2017Updated 9 years ago
- ☆17Oct 23, 2023Updated 2 years ago
- On Path to Multimodal Generalist: General-Level and General-Bench☆22Jul 11, 2025Updated last year
- An adaptive multispectral image fusion using particle swarm optimization☆15Dec 15, 2021Updated 4 years ago
- \infty-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation☆22Feb 14, 2025Updated last year
- Code for "Learning an adaptation function to assess image visual similarities", ICIP'21☆10Nov 12, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A basic implementation of the Bag of (Visual) Words approach (BoW) for image search.☆14Jul 9, 2014Updated 12 years ago
- [2025 TPAMI] Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation☆19Jan 3, 2026Updated 7 months ago
- This project explores the different techniques (both scalable and non scalable) for Graph based semi supervised learning. Recent techniqu…☆14May 28, 2016Updated 10 years ago
- [ICLR 2026] LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards.☆19Mar 16, 2026Updated 5 months ago
- Web Photo Source Identification based on Neural Enhanced Camera Fingerprint (WWW2023)☆15Feb 25, 2023Updated 3 years ago
- OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding☆24Apr 9, 2026Updated 4 months ago
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning☆57Jun 2, 2026Updated 3 months ago
- This is the Matlab code of paper "Fast Infrared and Visible Image Fusion with Structural Decomposition, Knowledge-Based Systems (KBS), 20…☆13Jun 22, 2020Updated 6 years ago
- Multi-focus image fusion using boosted random walks-based algorithm with two-scale focus maps☆12Jan 27, 2019Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆30Jun 26, 2026Updated 2 months ago
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆23Feb 21, 2026Updated 6 months ago
- The official implementation for the paper "Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything".☆24Nov 5, 2025Updated 9 months ago
- Scaling Test-time Training for LLM Reasoning☆34Apr 14, 2026Updated 4 months ago
- We introduce DreamPRM-1.5, an instance-reweighted framework that adaptively adjusts the importance of each training example via bi-level …☆16Nov 13, 2025Updated 9 months ago
- [ICCV 2025] A Benchmark for Multi-Step Reasoning in Long Narrative Videos☆28Jun 4, 2026Updated 2 months ago
- 3D-object detection using RGBD images☆16Jun 19, 2017Updated 9 years ago