Local self-attention in Transformer for visual question answering
☆13Mar 17, 2024Updated 2 years ago
Alternatives and similar repositories for LSAT
Users that are interested in LSAT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- opentqa is a open framework of the textbook question answering, which includes xtqa, mcan, cmr, mfb, mutan.☆11Mar 27, 2021Updated 5 years ago
- [ICCV 2021] Official implementation of the paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering"☆68Oct 11, 2021Updated 4 years ago
- ☆19May 31, 2023Updated 3 years ago
- Using image captions with LLM for zero-shot VQA☆19Mar 14, 2024Updated 2 years ago
- Repo for the EMNLP 2023 paper "A Simple Knowledge-Based Visual Question Answering"☆25Dec 14, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- RUArt: A Novel Text-Centered Solution for Text-Based Visual Question Answering☆10Nov 27, 2022Updated 3 years ago
- ☆15May 10, 2021Updated 5 years ago
- ☆10Oct 1, 2020Updated 5 years ago
- Official implementation of Dynamic Routing Transformer Network for Multimodal Sarcasm Detection (ACL'23)☆35Jul 9, 2023Updated 3 years ago
- the code for paper: A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question Answering☆14Aug 22, 2023Updated 2 years ago
- All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment☆21Feb 11, 2025Updated last year
- The official implementation of the paper "DIP: Dual Incongruity Perceiving Network for Sarcasm Detection"