MMA: Multimodal Memory Agent
☆23Mar 30, 2026Updated 3 months ago
Alternatives and similar repositories for MMA
Users that are interested in MMA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [preprint] Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning☆19Feb 18, 2026Updated 5 months ago
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆27Apr 17, 2026Updated 3 months ago
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning☆36Jul 1, 2026Updated 2 weeks ago
- Code and data for "Timo: Towards Better Temporal Reasoning for Language Models" (COLM 2024)☆26Oct 23, 2024Updated last year
- [FSE'2026] PlayCoder: Making LLM-Generated GUI Code Playable☆41Apr 22, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆11Dec 27, 2022Updated 3 years ago
- Cross-modal Coherence Modeling for Caption Generation☆11Jul 24, 2020Updated 5 years ago
- Open-Pandora: On-the-fly Control Video Generation☆35Nov 28, 2024Updated last year
- ☆47Apr 7, 2026Updated 3 months ago
- LatentMem: Customizing Latent Memory for Multi-Agent Systems☆48Feb 9, 2026Updated 5 months ago
- A pytorch implementation of our paper Image Captioning with Inherent Sentiment (ICME 2021 Oral).☆11Jul 18, 2022Updated 4 years ago
- [ACL-26 (main)] From Verbatim to Gist Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video A…☆39Apr 19, 2026Updated 3 months ago
- Unofficial PyTorch implementation of "Composing Good Shots by Exploiting Mutual Relations"☆15May 13, 2022Updated 4 years ago
- Unsupervised specificity-guided optimization of Image Captioning models to encourage meaningful diversity in the generated captions. Code…☆13May 25, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Generating high-quality image-pairs and training InstructPix2Pix with SDXL☆14Apr 9, 2024Updated 2 years ago
- A paper list of Weakly Supervised Object Detection (WSOD) resources.☆13May 6, 2021Updated 5 years ago
- [ICML 2026 Oral] Agent-native Mid-training for Software Engineering☆71Jun 7, 2026Updated last month
- Code for NAACL 2025 paper "AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge"☆16Mar 2, 2026Updated 4 months ago
- daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently☆38Feb 4, 2026Updated 5 months ago
- My tests and experiments with some popular dl frameworks.☆17Sep 11, 2025Updated 10 months ago
- Agentic System, Tool Use, Electronic Health Record, Large Language Models, Clinical Nature Language Processing☆24Apr 13, 2026Updated 3 months ago
- [EMNLP 2024] SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information☆11Oct 11, 2024Updated last year
- ☆15Apr 24, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation for "Mixture of In-Context Experts Enhance LLMs’ Awareness of Long Contexts" (Accepted by Neurips2024)☆14Jan 7, 2025Updated last year
- Official Project Page for Interactive Benchmarks☆31May 12, 2026Updated 2 months ago
- ControlLM is a method to control the personality traits and behaviors of language models in real-time at inference without costly trainin…☆21Nov 6, 2024Updated last year
- [ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents☆236May 13, 2026Updated 2 months ago
- Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning, release the dataset and the model weight☆13May 26, 2025Updated last year
- MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning☆68Jun 14, 2026Updated last month
- CAR-bench☆31Jul 13, 2026Updated last week
- Graph-based experience memory for LLM reward prediction with limited labels. 20% labels → 97.3% Oracle.☆19Mar 24, 2026Updated 3 months ago
- 中科大跨模态智能组-每周论文分享☆15Nov 20, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL 2026] WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering.☆15Apr 18, 2026Updated 3 months ago
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions☆15Jun 28, 2026Updated 3 weeks ago
- LLMs + Persona-Plug = Personalized LLMs☆15Oct 16, 2024Updated last year
- Look and Modify: Modification Networks for Image Captioning, BMVC 2019☆21Feb 18, 2020Updated 6 years ago
- ☆17Sep 23, 2024Updated last year
- Comparing sequential forecasters via confidence sequences & e-processes☆11Oct 24, 2023Updated 2 years ago
- [CVPR2025] BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding☆55Feb 5, 2026Updated 5 months ago