Official PyTorch codebase for the Modeling Caption Diversity in ContrastiveVision-Language Pretraining paper.
☆19Mar 28, 2025Updated last year
Alternatives and similar repositories for Llip
Users that are interested in Llip are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NAACL 2024] Z-GMOT: Zero-shot Generic Multiple Object Tracking☆12May 19, 2026Updated 4 months ago
- A curated list of Survey Papers on Deep Learning.☆13Sep 5, 2023Updated 3 years ago
- ☆15Jun 9, 2025Updated last year
- ☆18Mar 2, 2026Updated 6 months ago
- Code of "Robustifying Token Attention for Vision Transformers"☆20Dec 31, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Retrieval_OOD_for_Multimodal_AI☆11Dec 4, 2024Updated last year
- ☆18Nov 19, 2024Updated last year
- Official codes of the 1st place for The NVIDIA AI City Challenge 2023 - Track 2☆21Jul 25, 2023Updated 3 years ago
- Code accompanying paper "SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation"☆30May 8, 2026Updated 4 months ago
- [EMNLP 2024 Main] Official implementation of the paper "To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimoda…☆16Dec 13, 2024Updated last year
- 비디오 기반 인공지능 대화시스템☆11Aug 16, 2023Updated 3 years ago
- [ICCV 2023] Simple Baselines for Interactive Video Retrieval with Questions and Answers☆20Apr 16, 2024Updated 2 years ago
- Text-based Video Retrieval☆15Dec 4, 2024Updated last year
- ☆14Jan 5, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official repository for the article Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models (https://arxiv.org/…☆38Sep 5, 2025Updated last year
- ☆11May 1, 2023Updated 3 years ago
- [ICML 2026] Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions☆25Feb 11, 2026Updated 7 months ago
- Multimodal_AI_Video_Dialogue☆16Dec 3, 2024Updated last year
- ☆13Nov 7, 2021Updated 4 years ago
- ☆28Mar 13, 2025Updated last year
- In this project, facial recognition algorithm is implemented with python using PCA and SVD dimensionality reduction tools.☆11Sep 2, 2019Updated 7 years ago
- Hướng dẫn tạo một hệ thống Log Remote dùng chung cho nhiều dự án/server☆15Feb 26, 2020Updated 6 years ago
- My PhD manuscript LaTeX code and the slides for the defense☆11Feb 2, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This repository including most of cnn visualizations techniques using pytorch☆14Apr 14, 2020Updated 6 years ago
- ☆28Mar 13, 2025Updated last year
- All the codes of ACMICPC for two categories: Competition(Codeforces, Bestcoder, Regional, etc..) and Algorithms☆13Sep 23, 2017Updated 9 years ago
- This repository contains code for deploying a Gradio application using the SAM2 model for video processing. The application allows users …☆47Sep 24, 2024Updated 2 years ago
- ☆15May 7, 2024Updated 2 years ago
- Awesome-Text2Motion-Generation☆18Oct 26, 2023Updated 2 years ago
- [ECCV 2024] Official Release of SILC: Improving vision language pretraining with self-distillation☆48Oct 3, 2024Updated last year
- Official code for "SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation", AAAI 2024.☆37Jan 22, 2025Updated last year
- The official code for ICCV 2023 paper "Reconstructing Groups of People with Hypergraph Relational Reasoning"☆15Jul 4, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- F-16 is a powerful video large language model (LLM) that perceives high-frame-rate videos, which is developed by the Department of Electr…☆41Jul 3, 2025Updated last year
- Official Release of NeurIPS 2024 paper "Slot State Space Models"☆11Mar 22, 2025Updated last year
- A toolkit for computing Video Fréchet Inception Distance (VFID) metrics.☆11May 28, 2024Updated 2 years ago
- ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model☆15Apr 7, 2026Updated 5 months ago
- GPT-style network for phonemization with durations of text☆68Mar 21, 2024Updated 2 years ago
- Base repo for paper 'StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval'☆15Apr 27, 2022Updated 4 years ago
- use Blender software to visualize mesh sequences☆24Sep 2, 2019Updated 7 years ago