Official Repo for the paper: VCR: Visual Caption Restoration. Check arxiv.org/pdf/2406.06462 for details.
☆32Feb 26, 2025Updated last year
Alternatives and similar repositories for VCR
Users that are interested in VCR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Aug 9, 2026Updated last week
- ☆33Jul 3, 2025Updated last year
- ☆17Feb 22, 2024Updated 2 years ago
- ☆13May 9, 2023Updated 3 years ago
- Easy no-frills Jax implementations of common abstractions for simple diffusion models.☆11Feb 23, 2026Updated 5 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [SCIS 2024] The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Di…☆64Nov 7, 2024Updated last year
- [NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs☆160Apr 22, 2025Updated last year
- [CVPR 2024] LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation☆13Jun 17, 2024Updated 2 years ago
- MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering. A comprehensive evaluation of multimodal large model multilingua…☆64May 15, 2025Updated last year
- Variance Covariance Regularization☆14Jun 22, 2023Updated 3 years ago
- The code repository for "Wings: Learning Multimodal LLMs without Text-only Forgetting" [NeurIPS 2024]☆27Dec 28, 2024Updated last year
- ☆17Oct 22, 2024Updated last year
- [CVPR 2024] DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model☆19Apr 16, 2024Updated 2 years ago
- GroundCUA☆134Mar 24, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Linked Data to Natural Language☆11Jan 6, 2024Updated 2 years ago
- Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 30+ benchmarks☆15Feb 17, 2025Updated last year
- Official repository of MMDU dataset☆109Sep 29, 2024Updated last year
- [ICML 2024] | MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI☆119Apr 6, 2026Updated 4 months ago
- Code for Paper: Harnessing Webpage Uis For Text Rich Visual Understanding☆54Dec 12, 2024Updated last year
- ☆20Jan 10, 2025Updated last year
- CatMAE☆15Dec 13, 2023Updated 2 years ago
- PaCE: Parsimonious Concept Engineering for Large Language Models (NeurIPS 2024)☆43Jan 18, 2026Updated 7 months ago
- ☆47Nov 8, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Enable Comprehensive LLM Evaluation on Graph Reasoning☆81Jun 12, 2025Updated last year
- [ACL 2025] Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL☆16Oct 9, 2025Updated 10 months ago
- DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling☆38Jul 12, 2024Updated 2 years ago
- ☆16Nov 9, 2025Updated 9 months ago
- [ICLR 2025] Mathematical Visual Instruction Tuning for Multi-modal Large Language Models☆156Dec 5, 2024Updated last year
- ☆18Jun 12, 2024Updated 2 years ago
- Official repository for the paper "ModelTables: A Corpus of Tables about Models"☆16Updated this week
- Dataset Resplitting for Generalization in KGQA. See also https://github.com/semantic-systems/KGQA-datasets☆17Jun 29, 2022Updated 4 years ago
- Code and models for the paper "The effectiveness of MAE pre-pretraining for billion-scale pretraining" https://arxiv.org/abs/2303.13496☆93Mar 24, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Capstone Research Project in NYU Courant☆12Jan 3, 2020Updated 6 years ago
- Source code for "Taming GANs with Lookahead–Minmax", ICLR 2021.☆15Mar 28, 2021Updated 5 years ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- 非官方的MDCSpell论文的实现☆18Oct 16, 2022Updated 3 years ago
- Submissions, baselines and evaluations scripts for the 2nd version of the WebNLG+ Challenge 2020☆13Feb 1, 2022Updated 4 years ago
- The official implementation of the paper "MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding". …☆62Nov 5, 2024Updated last year
- ☆10Apr 13, 2020Updated 6 years ago