Reading list for Multimodal Large Language Models
☆70Aug 17, 2023Updated 3 years ago
Alternatives and similar repositories for Awesome-Multimodal-LLM
Users that are interested in Awesome-Multimodal-LLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Research Trends in LLM-guided Multimodal Learning.☆355Oct 17, 2023Updated 2 years ago
- Code and Data for the ACL21 paper "Modeling Bilingual Conversational Characteristics for Neural Chat Translation"☆13Dec 17, 2021Updated 4 years ago
- Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Mod…☆377Jul 3, 2026Updated 2 months ago
- [WIP@Oct 13] 质衡-基准测试 (Q-Bench in Chinese),包含中文版【底层视觉问答】和【底层视觉描述】数据集,以及中文提示下的图片质量评价。 We will release Q-Bench in more languages in the futu…☆24Jan 7, 2024Updated 2 years ago
- [AAAI 2023] Official repository of "Progressive Few-Shot Adaptation of Generative Model with Align-Free Spatial Correlation"☆10Jul 4, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria☆78Oct 16, 2024Updated last year
- An unofficial implementation of SOLAR-10.7B model and the newly proposed interlocked-DUS(iDUS) implementation and experiment details.☆14Mar 20, 2024Updated 2 years ago
- Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.☆761May 21, 2026Updated 3 months ago
- Implementation of the model: "(MC-ViT)" from the paper: "Memory Consolidation Enables Long-Context Video Understanding"☆27Aug 28, 2026Updated last week
- This is the official repository for the code and datasets in the paper "Progressive Open Space Expansion for Open-Set Model Attribution",…☆25Oct 22, 2023Updated 2 years ago
- Open Set Video HOI detection from Action-centric Chain-of-Look Prompting, ICCV2023☆12Oct 3, 2023Updated 2 years ago
- Recent Mexican Election Vote Returns☆12Updated this week
- Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval☆16Nov 29, 2025Updated 9 months ago
- ☆102Dec 22, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official code for paper "Reasoning Fails Where Step Flow Breaks" (ACL 2026)☆19Apr 19, 2026Updated 4 months ago
- 💬A curated list of incredible amount of publications related to Dialogue Systems especially Chatbots and Chit-chat Systems☆10Dec 5, 2019Updated 6 years ago
- Adversarial Category Alignment Network for Cross-domain Sentiment Classification (NAACL 2019)☆23Jul 4, 2019Updated 7 years ago
- TopicGPT allows to integrate the benefits of LLMs into Topic Modelling☆28Jun 22, 2024Updated 2 years ago
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot Classification☆11Aug 12, 2023Updated 3 years ago
- Pytorch version of VidLanKD: Improving Language Understanding viaVideo-Distilled Knowledge Transfer (NeurIPS 2021))☆56Feb 6, 2023Updated 3 years ago
- Gender prediction of chinese name based on LSTM☆14Mar 16, 2023Updated 3 years ago
- TyDiP Multilingual Politeness dataset and code☆12Oct 15, 2023Updated 2 years ago
- [ Arxiv 2023 ] This repository contains the code for "MUPPET: Multi-Modal Few-Shot Temporal Action Detection"☆16Aug 30, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Signal recovery and sampling over graphs☆16Oct 21, 2018Updated 7 years ago
- Scripts to evaluate various bias metrics for different NLG models + decoding algorithms☆16Dec 6, 2023Updated 2 years ago
- The source code of paper "CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking"☆81Dec 16, 2022Updated 3 years ago
- [EMNLP'23 Oral] ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain Dialogue PyTorch Implementation☆12Dec 4, 2023Updated 2 years ago
- Corpus analyses confrontation☆21Jan 24, 2023Updated 3 years ago
- ☆12Jul 8, 2019Updated 7 years ago
- Latest Advances on Multimodal Large Language Models☆18,002Updated this week
- ☆13Feb 7, 2023Updated 3 years ago
- XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts☆36Jul 2, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Model☆282Jun 25, 2024Updated 2 years ago
- Code for Chinese grammatical error correction based on knowledge distillation☆11Aug 16, 2022Updated 4 years ago
- A method for training neural networks that are provably robust to adversarial attacks. [IJCAI 2019]☆10Sep 3, 2019Updated 7 years ago
- [ICCV 2025] Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation☆29Aug 16, 2025Updated last year
- A Survey on multimodal learning research.☆332Aug 22, 2023Updated 3 years ago
- [ECCV 2024] Beyond MOT: Semantic Multi-Object Tracking☆31Sep 12, 2024Updated last year
- ☆14Aug 12, 2025Updated last year