[TMLR 2025] Reading List of Memory Augmented Multimodal Research, including multimodal context modeling, memory in vision and robotics, and external memory/knowledge augmented MLLM.
☆71Jan 17, 2026Updated 8 months ago
Alternatives and similar repositories for Awesome-Multimodal-Memory
Users that are interested in Awesome-Multimodal-Memory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PyTorch DataLoader for many VQA datasets☆15Jan 10, 2023Updated 3 years ago
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- This paper is currently under review by IEEE TCSVT, and the diffusion framework of the FedDiff algorithm part will be disclosed.☆14Mar 8, 2024Updated 2 years ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆20Apr 2, 2025Updated last year
- Curated systems, benchmarks, and papers etc. on memory for LLMs/MLLMs --- long-term context, retrieval, and reasoning.☆648Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ongoing project: a library for graph foundation model☆13Feb 7, 2024Updated 2 years ago
- Official repository for ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use☆31Nov 4, 2025Updated 10 months ago
- Official code for our EMNLP2021 Outstanding Paper MindCraft: Theory of Mind Modeling for Situated Dialogue in Collaborative Tasks☆21May 18, 2023Updated 3 years ago
- ☆13Oct 13, 2025Updated 11 months ago
- GitHub Repository for KDD 2022 paper "Saliency-Regularized Deep Multi-Task Learning"☆12Sep 26, 2023Updated 2 years ago
- Code and data for the paper Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization☆23Nov 21, 2022Updated 3 years ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆25Jun 26, 2026Updated 2 months ago
- Online video temporal grounding☆16Oct 20, 2025Updated 11 months ago
- DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric Voxelization (ICLR2024) & DynaVol-S: Dynamic Scene Understanding…☆21Apr 10, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The source code of PGD-VAE☆12Dec 8, 2022Updated 3 years ago
- Official code release for the NeurIPS 2021 article Training for the Future: A Simple Gradient Interpolation Loss to Generalize Along Time…☆10Nov 17, 2021Updated 4 years ago
- Uncertainty quantification for in-context learning of large language models☆15Apr 1, 2024Updated 2 years ago
- [NeurIPS 2024] Official Repository of Multi-Object Hallucination in Vision-Language Models☆37Nov 13, 2024Updated last year
- A controlled benchmark on evaluating and studying the dynamics of Long Context Language Models☆26Oct 17, 2025Updated 11 months ago
- QuoteSum is a textual QA dataset containing Semi-Extractive Multi-source Question Answering (SEMQA) examples written by humans, based on …☆13Mar 25, 2024Updated 2 years ago
- Code for EMNLP'24 paper - On Diversified Preferences of Large Language Model Alignment☆16Aug 6, 2024Updated 2 years ago
- ☆155Nov 17, 2025Updated 10 months ago
- (3DV 2026 Oral) L4P -- a feed-forward foundational model designed for multiple low-level 4D vision perception tasks.☆76Dec 9, 2025Updated 9 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A Deep Architecture for Synapse Detection in Multiplexed Fluorescence Images☆14May 28, 2019Updated 7 years ago
- [Paper][EMNLP 2025] RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models☆18Jan 29, 2026Updated 7 months ago
- ☆36May 24, 2025Updated last year
- Fast Spatial Memory with Elastic Test-Time Training (4D-LRM + 4D-LVSM)☆108Jun 20, 2026Updated 3 months ago
- [ACL 2026 oral] SeLaR: Selective Latent Reasoning in Large Language Models☆24Apr 25, 2026Updated 5 months ago
- ☆11May 24, 2024Updated 2 years ago
- repository for training action-conditioned latent diffusion world models for robot video generation☆80Aug 7, 2026Updated last month
- [EMNLP 2024] A Video Chat Agent with Temporal Prior☆34Mar 2, 2025Updated last year
- A curated list of awesome open-source libraries for context engineering (Long-term memory, MCP: Model Context Protocol, Prompt/RAG Compre…☆113Jul 3, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- PyTorch implementation of the NCDSSM models presented in the ICML '23 paper "Neural Continuous-Discrete State Space Models for Irregularl…☆27Jul 9, 2023Updated 3 years ago
- Code for AISTATS'25 paper - On the Power of Adaptive Weighted Aggregation in Heterogeneous Federated Learning and Beyond☆14Sep 23, 2025Updated last year
- ICCV 2025: Official Implematation of "Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced L…☆80Oct 25, 2025Updated 11 months ago
- 3DV 2024: Fast High Dynamic Range Radiance Fields for Dynamic Scenes☆34Aug 9, 2024Updated 2 years ago
- [CVPR 2025] 3D-GRAND: Towards Better Grounding and Less Hallucination for 3D-LLMs☆54Jun 13, 2024Updated 2 years ago
- ☆27Aug 6, 2026Updated last month
- NeurIPS 2021 paper 'Representation Learning on Spatial Networks' code☆18Oct 26, 2021Updated 4 years ago