☆97Feb 5, 2026Updated 8 months ago
Alternatives and similar repositories for VisMem
Users that are interested in VisMem are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Jul 31, 2025Updated last year
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement☆22Aug 4, 2026Updated 2 months ago
- Official repository for "Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models"☆24Dec 2, 2025Updated 10 months ago
- Official code repository for Med-CMR : "A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multi…☆28Dec 10, 2025Updated 10 months ago
- [ICLR 26] Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow☆47Oct 3, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [CVPR 2026] Boosting Reasoning in Large Multimodal Models via Activation Replay☆26Aug 25, 2026Updated last month
- [ICML 2026] The official code of FeRA: Frequency–Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning☆29Dec 27, 2025Updated 9 months ago
- ☆23May 26, 2025Updated last year
- From Large Angles to Consistent Faces: Identity-Preserving Video Generation via Mixture of Facial Experts☆27Jan 12, 2026Updated 8 months ago
- ☆29Nov 28, 2025Updated 10 months ago
- ☆187Jun 8, 2026Updated 4 months ago
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents☆424Sep 24, 2026Updated 2 weeks ago
- [ICML 2026] Transform Trained Transformer for Accelerating Native 4K Video Generation☆42Updated this week
- A paper list of Awesome Latent Space.☆977Jul 13, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"☆229Mar 19, 2026Updated 6 months ago
- [EMNLP'26 Findings] OPD-Evolver☆45Jun 17, 2026Updated 3 months ago
- ☆32Mar 22, 2026Updated 6 months ago
- [CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"☆88May 12, 2026Updated 4 months ago
- [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens☆300Aug 2, 2025Updated last year
- [ECCV 26'] Official codebase for the paper LaViT☆35Jul 30, 2026Updated 2 months ago
- [CVPR 2026] Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation☆64Dec 16, 2025Updated 9 months ago
- Official codebase for the paper Latent Visual Reasoning☆187Oct 22, 2025Updated 11 months ago
- ☆36Apr 22, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Awesome latest models, datasets and benchmarks on streaming/online video understanding.☆31Oct 19, 2025Updated 11 months ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism☆18Nov 18, 2025Updated 10 months ago
- [CVPR 2025 Oral] VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection☆143Jul 28, 2025Updated last year
- Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned percept…☆332Sep 27, 2026Updated last week
- [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?☆100Jul 13, 2025Updated last year
- [ICCV 2025] MMReason, MLLMs, step by step, reasoning benchmark, AGI☆15Apr 25, 2026Updated 5 months ago
- CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal☆26May 25, 2026Updated 4 months ago
- 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding☆80May 26, 2026Updated 4 months ago
- ☆88Jul 28, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆20Jun 10, 2025Updated last year
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆19Jul 15, 2025Updated last year
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- ☆26Oct 9, 2025Updated last year
- [AAAI 2026] SIFThinker: Spatially-Aware Image Focus for Visual Reasoning☆23Dec 2, 2025Updated 10 months ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆268Oct 18, 2025Updated 11 months ago
- ☆20Aug 7, 2025Updated last year