Feature extraction and visualization scripts for nocaps baselines.
☆18Jan 22, 2021Updated 5 years ago
Alternatives and similar repositories for image-feature-extractors
Users that are interested in image-feature-extractors are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Baseline model for nocaps benchmark, ICCV 2019 paper "nocaps: novel object captioning at scale".☆77Oct 3, 2023Updated 2 years ago
- Novel Object Captioner - Captioning Images with diverse objects☆42Nov 26, 2017Updated 8 years ago
- Website for TextVQA dataset.☆30Apr 30, 2023Updated 3 years ago
- ☆11Nov 13, 2020Updated 5 years ago
- Repository for the paper titled: "When is BERT Multilingual? Isolating Crucial Ingredients for Cross-lingual Transfer"☆13Nov 10, 2021Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for Decoupled Novel Object Captioner☆28Feb 26, 2020Updated 6 years ago
- Official code for the paper "Contrast and Classify: Training Robust VQA Models" published at ICCV, 2021☆19Jul 27, 2021Updated 5 years ago
- Pytorch version of VidLanKD: Improving Language Understanding viaVideo-Distilled Knowledge Transfer (NeurIPS 2021))☆56Feb 6, 2023Updated 3 years ago
- Multitask Multilingual Multimodal Pre-training☆72Nov 27, 2022Updated 3 years ago
- Dataset and models for paper "Game-Based Video-Context Dialogue (EMNLP 2018)"☆19Oct 25, 2018Updated 7 years ago
- ☆16Apr 27, 2021Updated 5 years ago
- GLAC Net: GLocal Attention Cascading Network for the Visual Storytelling Challenge☆45Aug 26, 2020Updated 5 years ago
- Implementation for "Large-scale Pretraining for Visual Dialog" https://arxiv.org/abs/1912.02379☆95Mar 31, 2020Updated 6 years ago
- ☆53Dec 13, 2019Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for WACV 2021 Paper "Meta Module Network for Compositional Visual Reasoning"☆43May 13, 2021Updated 5 years ago
- Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions. CVPR 2019☆281Dec 21, 2022Updated 3 years ago
- This is a python library. Install with "python3 -m pip install rp" then run with "python3 -m rp" or just "rp". Requires python≥3.5☆13Jul 13, 2026Updated 2 weeks ago
- Bottom-up attention model for image captioning and VQA, based on Faster R-CNN and Visual Genome☆23Aug 22, 2019Updated 6 years ago
- Show, Edit and Tell: A Framework for Editing Image Captions, CVPR 2020☆82Jul 17, 2020Updated 6 years ago
- (ECCV2024) Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence☆20Jun 27, 2025Updated last year
- 针对 markdown 文件的命令行翻译☆14Feb 2, 2023Updated 3 years ago
- Code for ICML 2019 paper "Probabilistic Neural-symbolic Models for Interpretable Visual Question Answering" [long-oral]☆68Aug 3, 2023Updated 2 years ago
- The source code and the data for ACL 2022 paper "Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web Dat…☆14Apr 21, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Stack-Captioning: Coarse-to-Fine Learning for Image Captioning☆63Apr 18, 2018Updated 8 years ago
- ☆12Dec 13, 2022Updated 3 years ago
- ☆40May 28, 2018Updated 8 years ago
- [TMI 2024] Harvard Glaucoma Fairness (Harvard-GF): A Retinal Nerve Disease Dataset for Fairness Learning and Fair Identity Normalization☆10Apr 9, 2024Updated 2 years ago
- Unpaired Image Captioning☆36Mar 25, 2021Updated 5 years ago
- Code for the CVPR 2020 paper 'Action Modifiers: Learning from Adverbs in Instructional Videos'☆23May 17, 2021Updated 5 years ago
- Meta-Reinforced Synthetic Data for One-Shot Fine-Grained Visual Recognition (NeurIPS 2019 & PAMI 2022)☆32Jul 10, 2024Updated 2 years ago
- Code base for the paper "Latent variable model for multi-modal translation".☆17Jul 25, 2024Updated 2 years ago
- ☆29Dec 8, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for the paper "Multi-Task Learning of Object States and State-Modifying Actions from Web Videos" published in TPAMI☆11Mar 3, 2024Updated 2 years ago
- ☆45Oct 11, 2021Updated 4 years ago
- Pytorch implementation of https://arxiv.org/pdf/1909.10470.pdf☆32Aug 23, 2021Updated 4 years ago
- Generalization in Metric Learning: Should the Embedding Layer be the Embedding Layer?☆11Jan 3, 2019Updated 7 years ago
- Code for the paper "Refining Language Model with Compositional Explanation" (NeurIPS 2021)☆11Oct 25, 2021Updated 4 years ago
- Transfer Learning for Named Entity Recognition☆10Mar 14, 2019Updated 7 years ago
- Simply implement the RetinaNet with BiFPN (EfficientDet) by tensorflow☆10Jun 15, 2020Updated 6 years ago