LAVIS - A One-stop Library for Language-Vision Intelligence
☆48Aug 5, 2024Updated 2 years ago
Alternatives and similar repositories for LAVIS-XInstructBLIP
Users that are interested in LAVIS-XInstructBLIP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17May 14, 2025Updated last year
- Official implementation of T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition☆20Oct 23, 2024Updated last year
- ☆21Jan 22, 2026Updated 6 months ago
- Creative Instructions Project☆11Sep 4, 2023Updated 2 years ago
- Visual Programming for Text-to-Image Generation and Evaluation (NeurIPS 2023)☆57Jul 25, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [EMNLP 2025 Main] The official repo of MMLU-ProX benchmark.☆29Aug 26, 2025Updated 11 months ago
- ☆44Feb 21, 2023Updated 3 years ago
- [CVPR 2024] OneLLM: One Framework to Align All Modalities with Language☆666Oct 22, 2024Updated last year
- Project website for "Telling left from right: Learning spatial correspondence between sight and sound"☆29Jun 6, 2022Updated 4 years ago
- Florence-2☆72Feb 13, 2025Updated last year
- [ICLR 2025] MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs☆56Mar 25, 2025Updated last year
- finetune your florence2 model easy☆21Jul 27, 2024Updated 2 years ago
- Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning☆44Jul 2, 2025Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official implementation of "Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data" (ICLR 2024)☆36Oct 16, 2024Updated last year
- Official This-Is-My Dataset published in CVPR 2023☆16Jul 18, 2024Updated 2 years ago
- ☆14May 21, 2024Updated 2 years ago
- [ISPRS 2024] LoveNAS: Towards Multi-Scene Land-Cover Mapping via Hierarchical Searching Adaptive Network☆33Dec 1, 2024Updated last year
- code for paper "Detecting Adversarial Data via Perturbation Forgery"☆18Jul 10, 2025Updated last year
- BEAR: a new BEnchmark on video Action Recognition☆46Apr 21, 2024Updated 2 years ago
- [ACM MM 2024] FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMs☆60Aug 8, 2024Updated 2 years ago
- This package offers core functionalities to retrieve all versions of the World Imagery Wayback archive, or find versions that contain loc…☆10Jun 10, 2026Updated 2 months ago
- Visual-Text dataset based on NFT metadata☆19Nov 7, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [TPAMI 2026] A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications☆16Jun 9, 2026Updated 2 months ago
- A curated list of 'Talking Head Generation' resources. Features influential papers, groundbreaking algorithms, crucial GitHub repositorie…☆75Oct 17, 2023Updated 2 years ago
- Framework for computationally efficient training of universal image feature extraction models.☆21Aug 19, 2024Updated last year
- ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning. In ICCV, 2021.☆64Nov 18, 2021Updated 4 years ago
- [EMNLP 2025 Findings] MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation☆15Aug 22, 2025Updated 11 months ago
- Resources for our IJCAI 2020 paper, TopicKA: Generating Commonsense Knowledge-Aware Dialogue Responses Towards the Recommended Topic Fact☆12Nov 30, 2020Updated 5 years ago
- [ICLR 2025] IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model☆37Nov 27, 2024Updated last year
- ☆11Oct 13, 2024Updated last year
- Java web application backed by the Ethereum-Blockchain network. Powered by RESTful web services (JAX-RS && Spring Boot) , Docker, Kuberne…☆15Feb 19, 2019Updated 7 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This is the official implementation of "GvSeg: General and Task-Oriented Video Segmentation" (Accepted at ECCV 2024).☆18Jul 15, 2024Updated 2 years ago
- ☆63Sep 23, 2024Updated last year
- ☆46Jun 23, 2026Updated last month
- Wasserstein Divergence for GANs☆19Jan 21, 2021Updated 5 years ago
- (ICCV23 Oral) LOGICSEG: Parsing Visual Semantics with Neural Logic Learning and Reasoning☆25Apr 11, 2024Updated 2 years ago
- ☆12Aug 7, 2024Updated 2 years ago
- Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise☆42Oct 7, 2025Updated 10 months ago