[TMLR 2025] Reading List of Memory Augmented Multimodal Research, including multimodal context modeling, memory in vision and robotics, and external memory/knowledge augmented MLLM.
☆69Jan 17, 2026Updated 6 months ago
Alternatives and similar repositories for Awesome-Multimodal-Memory
Users that are interested in Awesome-Multimodal-Memory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PyTorch DataLoader for many VQA datasets☆15Jan 10, 2023Updated 3 years ago
- Official implementation of Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time Training.☆40Apr 9, 2026Updated 3 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆18Apr 2, 2025Updated last year
- Curated systems, benchmarks, and papers etc. on memory for LLMs/MLLMs --- long-term context, retrieval, and reasoning.☆549Updated this week
- Ongoing project: a library for graph foundation model☆13Feb 7, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official repository for ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use☆31Nov 4, 2025Updated 8 months ago
- Code and dataset for paper: Multi-stage Deep Classifier Cascades for OpenWorld Recognition☆14Mar 20, 2020Updated 6 years ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆23Jul 14, 2026Updated last week
- Official code repository for the main conference paper in EMNLP 2022: SubeventWriter: Iterative Sub-event Sequence Generation with Cohere…☆11Oct 16, 2022Updated 3 years ago
- Official code for our EMNLP2021 Outstanding Paper MindCraft: Theory of Mind Modeling for Situated Dialogue in Collaborative Tasks☆21May 18, 2023Updated 3 years ago
- ☆13Oct 13, 2025Updated 9 months ago
- Code and data for the paper Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization☆23Nov 21, 2022Updated 3 years ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆24Jun 26, 2026Updated 3 weeks ago
- DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric Voxelization (ICLR2024) & DynaVol-S: Dynamic Scene Understanding…☆21Apr 10, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Online video temporal grounding☆16Oct 20, 2025Updated 9 months ago
- Official code release for the NeurIPS 2021 article Training for the Future: A Simple Gradient Interpolation Loss to Generalize Along Time…☆10Nov 17, 2021Updated 4 years ago
- Uncertainty quantification for in-context learning of large language models☆15Apr 1, 2024Updated 2 years ago
- [NeurIPS 2024] Official Repository of Multi-Object Hallucination in Vision-Language Models☆37Nov 13, 2024Updated last year
- A controlled benchmark on evaluating and studying the dynamics of Long Context Language Models☆26Oct 17, 2025Updated 9 months ago
- QuoteSum is a textual QA dataset containing Semi-Extractive Multi-source Question Answering (SEMQA) examples written by humans, based on …☆13Mar 25, 2024Updated 2 years ago
- Code for EMNLP'24 paper - On Diversified Preferences of Large Language Model Alignment☆16Aug 6, 2024Updated last year
- ☆155Nov 17, 2025Updated 8 months ago
- (3DV 2026 Oral) L4P -- a feed-forward foundational model designed for multiple low-level 4D vision perception tasks.☆72Dec 9, 2025Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A Deep Architecture for Synapse Detection in Multiplexed Fluorescence Images☆14May 28, 2019Updated 7 years ago
- [Paper][EMNLP 2025] RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models☆17Jan 29, 2026Updated 5 months ago
- ☆36May 24, 2025Updated last year
- Fast Spatial Memory with Elastic Test-Time Training (4D-LRM + 4D-LVSM)☆102Jun 20, 2026Updated last month
- ☆34Apr 4, 2024Updated 2 years ago
- [ACL 2026 oral] SeLaR: Selective Latent Reasoning in Large Language Models☆20Apr 25, 2026Updated 2 months ago
- ☆11May 24, 2024Updated 2 years ago
- Github Repo for ICML 2022 paper: Communication-Efficient Adaptive Federated Learning☆10Nov 18, 2022Updated 3 years ago
- repository for training action-conditioned latent diffusion world models for robot video generation☆73May 29, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- PyTorch implementation of the NCDSSM models presented in the ICML '23 paper "Neural Continuous-Discrete State Space Models for Irregularl…☆27Jul 9, 2023Updated 3 years ago
- Code for AISTATS'25 paper - On the Power of Adaptive Weighted Aggregation in Heterogeneous Federated Learning and Beyond☆14Sep 23, 2025Updated 9 months ago
- The implementation of paper ''Efficient Attention Network: Accelerate Attention by Searching Where to Plug''.☆20Jun 16, 2023Updated 3 years ago
- ICCV 2025: Official Implematation of "Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced L…☆76Oct 25, 2025Updated 8 months ago
- 3DV 2024: Fast High Dynamic Range Radiance Fields for Dynamic Scenes☆34Aug 9, 2024Updated last year
- [CVPR 2025] 3D-GRAND: Towards Better Grounding and Less Hallucination for 3D-LLMs☆54Jun 13, 2024Updated 2 years ago
- Implementation of "Group-Wise Deep Object Co-Segmentation With Co-Attention Recurrent Neural Network" ICCV 2019☆14Jan 27, 2023Updated 3 years ago