Lightweight Transformer for Multi-modal Tasks
☆16Dec 9, 2022Updated 3 years ago
Alternatives and similar repositories for LWTransformer
Users that are interested in LWTransformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SOIT: Segmenting Objects with Instance-Aware Transformers☆14Jun 6, 2022Updated 4 years ago
- Official implementation for "SimA: Simple Softmax-free Attention for Vision Transformers"☆48Apr 18, 2024Updated 2 years ago
- Optimized code based on M2 for faster image captioning training☆21Nov 18, 2022Updated 3 years ago
- CVPR2022 - Language-Bridged Spatial-Temporal Interaction for Referring Video Object Segmentation☆24Aug 12, 2022Updated 4 years ago
- Code for "Searching for Efficient Multi-Stage Vision Transformers"☆63Sep 1, 2021Updated 5 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Implementation of the paper ''Implicit Feature Refinement for Instance Segmentation''.☆20Oct 27, 2021Updated 4 years ago
- ☆21Feb 3, 2025Updated last year
- ☆13Oct 30, 2023Updated 2 years ago
- Deep Multimodal Neural Architecture Search☆29Nov 15, 2020Updated 5 years ago
- Official Implementation for paper "Referring Transformer: A One-step Approach to Multi-task Visual Grounding" Neurips 2021☆67May 26, 2022Updated 4 years ago
- Scene Graph Generate Zero Shot☆23Apr 16, 2023Updated 3 years ago
- INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model☆42Aug 4, 2024Updated 2 years ago
- ☆24Jun 27, 2025Updated last year
- ☆22Jun 30, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆35Jul 4, 2024Updated 2 years ago
- Using tensorflow object detection api and openCV to calculate real world coordinates from top view with fixed height of the camera.☆10Jun 19, 2021Updated 5 years ago
- Official Code for "Knowing what it is: Semantic-enhanced Dual Attention Transformer" (TMM2022)☆19Oct 15, 2022Updated 3 years ago
- PyTorch implementation of "Dynamic Structure Pruning for Compressing CNNs" (AAAI 2023 Oral)☆28Jan 15, 2024Updated 2 years ago
- The official implementation for SETA (TIP 2024).☆12Feb 17, 2025Updated last year
- (ACM MM24) This is the offical repository of GIST: Improving Parameter Efficient Fine Tuning via Knowledge Interaction.☆10Jan 28, 2024Updated 2 years ago
- Exploring Lightweight Structures for Tiny Object Detection in Remote Sensing Images (IEEE TGRS 2025)☆31Feb 10, 2026Updated 8 months ago
- A server/client approach to face recognition. Aims to be fast, secure and iot friendly. Uses dlib.☆11May 7, 2021Updated 5 years ago
- [ECCV 2022] AMixer: Adaptive Weight Mixing for Self-attention Free Vision Transformers☆29Nov 14, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆27May 22, 2024Updated 2 years ago
- Transformer and Neural Operator for solving Stochastic PDE☆12May 22, 2022Updated 4 years ago
- video captioning☆25Mar 14, 2019Updated 7 years ago
- [CVPR2022 Oral] 3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds☆57Jan 29, 2023Updated 3 years ago
- Measure the diversity of image descriptions, repository for our COLING 2018 paper.☆13Dec 29, 2019Updated 6 years ago
- This is the repo for "Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition", CVPR2025.☆30Dec 22, 2025Updated 9 months ago
- DOneLogin Android: Facial verification for Two-Factors Authentication (2FA) on Android platform☆11Mar 30, 2021Updated 5 years ago
- LV-BERT: Exploiting Layer Variety for BERT (Findings of ACL 2021)☆19May 10, 2023Updated 3 years ago
- GQA-OOD is a new dataset and benchmark for the evaluation of VQA models in OOD (out of distribution) settings.☆33Mar 1, 2021Updated 5 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [Codes of paper]: Busy-Quiet Video Disentangling for Video Classification☆14Jan 17, 2022Updated 4 years ago
- Partially Non-Autoregressive Image Captioning☆10Sep 30, 2021Updated 5 years ago
- A system for extracting surfaces from a Point Cloud using ROS and PCL☆10Oct 7, 2016Updated 10 years ago
- ☆10Apr 20, 2018Updated 8 years ago
- Code for the paper "Spot What Matters: Learning Context Using Graph Convolutional Networks for Weakly-Supervised Action Detection"☆14Aug 18, 2021Updated 5 years ago
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning, CVPR 2022☆98Dec 2, 2022Updated 3 years ago
- nocaps: novel object captioning at scale☆10May 23, 2019Updated 7 years ago