☆27Jul 20, 2024Updated 2 years ago
Alternatives and similar repositories for ConTextual
Users that are interested in ConTextual are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Jul 2, 2025Updated last year
- Code for paper "Point and Ask: Incorporating Pointing into Visual Question Answering"☆19Oct 4, 2022Updated 4 years ago
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- Code for paper: "Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines"☆10Oct 11, 2024Updated 2 years ago
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuning (ICML 2023)☆19Dec 15, 2023Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [NeurIPS 2024] Calibrated Self-Rewarding Vision Language Models☆87Oct 26, 2025Updated 11 months ago
- Official implementation of the TransT-M (the winner of VOT-RT 2021) , including code and models.☆28Mar 28, 2023Updated 3 years ago
- Learning Low-rank and Sparse Discriminative Correlation Filters for Coarse-to-Fine Visual Object Tracking☆10Apr 15, 2021Updated 5 years ago
- ☆19Sep 1, 2025Updated last year
- ☆29Jan 23, 2024Updated 2 years ago
- OpenVLThinker [NeurIPS 2025] & OpenVLThinkerV2 [COLM 2026]☆158May 25, 2026Updated 4 months ago
- (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.☆367Jan 14, 2025Updated last year
- The official repo of paper "Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller"☆18Aug 13, 2024Updated 2 years ago
- Official implementation of "What does CLIP know about a red circle? Visual Prompt Engineering for VLMs", ICCV 2023☆12Sep 21, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆27Aug 28, 2023Updated 3 years ago
- Code for 'Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality', EMNLP 2022☆31May 29, 2023Updated 3 years ago
- [ECCV2024, Oral, Best Paper Finalist] This is the official implementation of the paper "LEGO: Learning EGOcentric Action Frame Generation…☆41Feb 24, 2025Updated last year
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement☆21Jul 4, 2026Updated 3 months ago
- ☆90Sep 15, 2026Updated 3 weeks ago
- ☆11May 24, 2024Updated 2 years ago
- MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts☆367Jul 27, 2026Updated 2 months ago
- ☆51Oct 29, 2023Updated 2 years ago
- ☆18Dec 2, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [NAACL 2024] MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning☆95Jan 7, 2025Updated last year
- Explaining Deep Convolutional Neural Networks via Unsupervised Visual-Semantic Filter Attention (CVPR 2022)☆20Mar 31, 2022Updated 4 years ago
- An official implementation for "Global Tracking via Ensemble of Local Trackers"☆11Mar 13, 2022Updated 4 years ago
- ☆32Feb 8, 2024Updated 2 years ago
- Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"☆68Oct 19, 2024Updated last year
- Evaluation toolkit of the informative tracking benchmark comprising 9 scenarios, 180 diverse videos, and new challenges.☆17Dec 14, 2021Updated 4 years ago
- Not All Poisons are Created Equal: Robust Training against Data Poisoning (ICML 2022)☆23Aug 8, 2022Updated 4 years ago
- On the Hidden Mystery of OCR in Large Multimodal Models (OCRBench)☆896Updated this week
- ☆17Dec 22, 2021Updated 4 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- EMNLP2023 - InfoSeek: A New VQA Benchmark focus on Visual Info-Seeking Questions☆27May 30, 2024Updated 2 years ago
- A spoken version of the textual story cloze benchmark☆22Aug 6, 2023Updated 3 years ago
- Repo for paper: "Paxion: Patching Action Knowledge in Video-Language Foundation Models" Neurips 23 Spotlight☆38May 23, 2023Updated 3 years ago
- CaMML:Context-Aware MultiModal Learner for Large Models (ACL 2024 SAC Award)☆15May 21, 2025Updated last year
- How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?☆12Aug 16, 2023Updated 3 years ago
- ☆51Jun 14, 2024Updated 2 years ago
- [NLPCC'23] ZeroGen: Zero-shot Multimodal Controllable Text Generation with Multiple Oracles PyTorch Implementation☆14Oct 7, 2023Updated 3 years ago