Open source implementation of the paper "MM-Vid: Advancing Video Understanding with GPT-4V(ision)".
☆44Jan 4, 2026Updated 7 months ago
Alternatives and similar repositories for MM-VID
Users that are interested in MM-VID are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [AAAI2025] Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark☆31Aug 1, 2026Updated last week
- [CVPR2025] Number it: Temporal Grounding Videos like Flipping Manga☆150Aug 1, 2026Updated last week
- ☆20Sep 19, 2023Updated 2 years ago
- [NeurIPS 2024] Visual Perception by Large Language Model’s Weights☆56Mar 31, 2025Updated last year
- Auto-Rubric as Reward: From Implicit Preference to Explicit Generative Criteria☆51Jul 24, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An in-context learning research testbed☆19Mar 16, 2025Updated last year
- Using distilled CLIP model to deploy the android device☆20Feb 28, 2023Updated 3 years ago
- (NeurIPS 2024 Spotlight) TOPA: Extend Large Language Models for Video Understanding via Text-Only Pre-Alignment☆29Sep 27, 2024Updated last year
- [ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.☆1,095Aug 1, 2026Updated last week
- Training Vision Transformers for Semi-Supervised Semantic Segmentation☆16Nov 3, 2025Updated 9 months ago
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models (ICLR2026)☆23Jun 24, 2026Updated last month
- TrackGPT: Track What You Need in Videos via Text Prompts☆25May 16, 2023Updated 3 years ago
- [ICML 2025] Official code of "DAMA: Data- and Model-aware Alignment of Multi-modal LLMs"☆16May 24, 2025Updated last year
- [TPAMI2025] BackMix: Regularizing Open Set Recognition by Removing Underlying Fore-Background Priors☆16Apr 23, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆11May 24, 2024Updated 2 years ago
- A mobile GUI search engine using a vision-language model☆15May 5, 2025Updated last year
- Remon-OBS-Studio(ROS): Program that broadcasts using WebRTC and is based on obs-studio proejct and Pion WebRTC project.☆11Sep 6, 2019Updated 6 years ago
- Personalized Image Generation with Large Multimodal Models☆17May 13, 2025Updated last year
- A curated list of all awesome pygames created by Agneay B Nair☆11Apr 28, 2024Updated 2 years ago
- Svelte and tailwindcss slider (range input)☆10Apr 10, 2021Updated 5 years ago
- A dataset with classified film shots☆11Aug 8, 2022Updated 4 years ago
- [ICCV2023] CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection☆19Apr 23, 2025Updated last year
- Demo☆13Jan 7, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Performant video filters in a browser using webassembly. Implements various video scopes such as lumascope, rgb parade, and vector scope.☆10Jun 3, 2021Updated 5 years ago
- Azure OpenAI benchmarking tool☆29Apr 4, 2025Updated last year
- Video Chain of Thought, Codes for ICML 2024 paper: "Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition"☆182Feb 25, 2025Updated last year
- Tools to create and distribute macOS Applications through disk images☆17Apr 18, 2020Updated 6 years ago
- ☆11Jun 11, 2024Updated 2 years ago
- ☆57Nov 21, 2024Updated last year
- Official implementation of "TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards"☆32Jul 20, 2026Updated 3 weeks ago
- PyTorch code for "ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning"☆21Oct 28, 2024Updated last year
- NeurIPS 2024: SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation☆13May 24, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- 「ECCV 2024」 PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation☆21Jul 2, 2024Updated 2 years ago
- The eWeather is a GUI application that allows you to get the weather forecast for any city in the world. It is written in C++ and Qt and …☆15Dec 2, 2022Updated 3 years ago
- Baidu Qianfan Deep Research☆36Jun 8, 2026Updated 2 months ago
- EARL: Editing with Autoregression and RL☆43Nov 21, 2025Updated 8 months ago
- 一个小小的书单,收集整理了一些计算机科学与技术方面的书籍英文原著pdf。☆10Jan 13, 2022Updated 4 years ago
- Details and sample code for the Eval Engineering course☆16Jan 27, 2026Updated 6 months ago
- ⚡️Qwen-Image 4.8x🎉 speedup with Hybrid Acceleration for low VRAM GPUs☆17Oct 24, 2025Updated 9 months ago