Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
☆16Nov 29, 2025Updated 8 months ago
Alternatives and similar repositories for RGE
Users that are interested in RGE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML2026] FreeRet: MLLMs as Training-Free Retrievers☆24May 25, 2026Updated 2 months ago
- A Fine-grained Benchmark for Video Captioning and Retrieval☆30Jul 16, 2025Updated last year
- [ECCV 2024 Oral] SPLAM: Accelerating Image Generation with Sub-path Linear Approximation Model☆24Nov 1, 2024Updated last year
- A simple Computer Vision Framework, mainly based on PyTorch. Including distributed training, logging and so on.☆12Dec 2, 2023Updated 2 years ago
- ☆46Jun 23, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The code implementation for UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings (ICLR 2026).☆71Feb 25, 2026Updated 5 months ago
- A visualization tool for temporal action localization (detection/segmentation).☆13Mar 30, 2023Updated 3 years ago
- VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model☆15Jul 31, 2025Updated last year
- [ICLR2023] Video Scene Graph Generation from Single-Frame Weak Supervision☆12Sep 17, 2023Updated 2 years ago
- [ICML 2025] Differentiable Solver Search for Fast Diffusion Sampling☆21Jul 7, 2025Updated last year
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsing☆28Jan 19, 2026Updated 6 months ago
- [ECCV 2022] Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing☆27Jul 15, 2022Updated 4 years ago
- ☆13Jun 11, 2026Updated last month
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval (ICCV 2025 Highlight)☆27Aug 1, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- An official implementation for MS-DETR in ACL'23☆17Jun 3, 2023Updated 3 years ago
- Region Encoder Network☆20Oct 2, 2025Updated 10 months ago
- ☆11May 10, 2024Updated 2 years ago
- Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization☆16Jul 20, 2023Updated 3 years ago
- [ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models☆20Jul 17, 2024Updated 2 years ago
- Pytorch Implementation of ECCV'22 paper: Video Activity Localisation with Uncertainties in Temporal Boundary☆17Jul 17, 2022Updated 4 years ago
- [CVPR 2025] DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval☆22Jun 23, 2025Updated last year
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- The Source Code for IF-VidCap @ICLR 2026☆19Oct 22, 2025Updated 9 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs☆117Jul 27, 2026Updated last week
- [CVPR 2023 Hightlight] PDPP: Projected Diffusion for Procedure Planning in Instructional Videos☆34Aug 30, 2023Updated 2 years ago
- [CVPR'25] Official implementation of the paper "Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Mo…☆18Nov 21, 2025Updated 8 months ago
- Source code of the paper Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval☆19May 13, 2026Updated 2 months ago
- [ECCV 2026] History-Aware Transformation of ReID Features for Multiple Object Tracking☆37Jul 23, 2026Updated 2 weeks ago
- Official Implementation of ISR-DPO:Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO (AAAI'25)☆23Nov 25, 2025Updated 8 months ago
- [ICCV 2023] The official PyTorch implementation of the paper: "Localizing Moments in Long Video Via Multimodal Guidance"☆23Sep 26, 2024Updated last year
- Composed Video Retrieval☆62May 2, 2024Updated 2 years ago
- [CVPR'25 Highlight] Official implementation for paper - LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis☆161Apr 15, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Open Set Video HOI detection from Action-centric Chain-of-Look Prompting, ICCV2023☆12Oct 3, 2023Updated 2 years ago
- [IJCV] Progressive Visual Prompt Learning with Contrastive Feature Re-formation☆15Aug 10, 2024Updated 2 years ago
- Pytorch implementation of the paper 'Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Super…☆19Jan 19, 2024Updated 2 years ago
- [ECCV 2024] ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video☆23Jul 29, 2024Updated 2 years ago
- Motion-Aware Generative Frame Interpolation☆50Mar 11, 2025Updated last year
- [ICCV 2023] Robust Object Modeling for Visual Tracking, Official Implementation☆48Jan 5, 2025Updated last year
- a pytorch implement of Supervised Contrastive Learning with memory bank(queue)☆15May 25, 2022Updated 4 years ago