This is a repository for awesome any2any work collection.
☆30Jul 10, 2026Updated 3 weeks ago
Alternatives and similar repositories for awesome-any2any
Users that are interested in awesome-any2any are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- On Path to Multimodal Generalist: General-Level and General-Bench☆21Jul 11, 2025Updated last year
- MMDeepResearch-Bench (MMDR)☆32Apr 1, 2026Updated 4 months ago
- [AAAI 2026] Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing☆24Nov 20, 2025Updated 8 months ago
- the official code of "Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation" (ECCV2024)☆13Jan 14, 2025Updated last year
- (ICML 2026) Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search☆50May 1, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- The official repository of paper "Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark"☆19Jun 20, 2025Updated last year
- (ICCV 2021) Official PyTorch implementation of "Learning to Discover Reflection Symmetry via Polar Matching Convolution."☆13Aug 31, 2021Updated 4 years ago
- [WACV 2025] Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge☆41Oct 29, 2024Updated last year
- Reproduction code for paper "MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft"☆18Jun 12, 2026Updated last month
- [ECCV 2026] Official Implementation of Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction☆18Apr 26, 2026Updated 3 months ago
- Official implementation for the AAAI2025 paper "PIXELS - Progressive Image Xemplar-based Editing with Latent Surgery"☆11Dec 17, 2024Updated last year
- [AAAI26] Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilitie…☆11Feb 7, 2026Updated 5 months ago
- Official implementation of AAAI-2024 paper "Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain"☆13Jun 17, 2024Updated 2 years ago
- Code of the Grounded MUIE model, REAMO☆11Dec 3, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations☆22Jun 17, 2026Updated last month
- ☆20Sep 1, 2025Updated 11 months ago
- Code release for "MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning"☆11Oct 11, 2024Updated last year
- [CVPR'25] 🌟🌟 EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering☆52Jun 19, 2025Updated last year
- Implementation of NAACL'19 Strong and Simple Baselines for Multimodal Utterance Embeddings☆10Jun 4, 2019Updated 7 years ago
- [ICLR'25 Oral] MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models☆35Nov 3, 2024Updated last year
- The repository of paper Personalized Multimodal Response Generation with Large Language Models☆18Jun 28, 2024Updated 2 years ago
- ☆12Feb 17, 2025Updated last year
- a py3 lib for NLP & image-caption metrics : BLEU METEOR CIDEr ROUGE SPICE WMD☆14Sep 13, 2022Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Awesome Unified Multimodal Models☆1,310Mar 24, 2026Updated 4 months ago
- We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their…☆22Jan 11, 2026Updated 6 months ago
- ☆15Dec 13, 2022Updated 3 years ago
- A PyTorch implementation of "VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics"☆19Mar 9, 2026Updated 4 months ago
- ☆13Jan 18, 2024Updated 2 years ago
- Implementation of the Paper Scene-Graph ViT☆10Dec 20, 2024Updated last year
- A lightweight, agent-style framework for fact-checking atomic claims using iterative retrieval and verification. Reduces LLM and search c…☆20Jun 4, 2025Updated last year
- Codex skill for reconstructing diagram images into editable Draw.io files☆21Jul 21, 2026Updated 2 weeks ago
- [ICLR 26] The official code repository for the paper "Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions".☆18Feb 9, 2026Updated 5 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- The implementation of paper "Self-supervised learning for multimedia recommendation", TMM'22.☆11Jul 4, 2022Updated 4 years ago
- NeurIPS'23: Energy Discrepancies: A Score-Independent Loss for Energy-Based Models☆18Oct 22, 2024Updated last year
- ☆12Feb 2, 2024Updated 2 years ago
- ☆11Aug 10, 2024Updated last year
- ☆17Mar 2, 2023Updated 3 years ago
- VibeQuant: Your Personal Quant Research Workbench☆18Jul 9, 2026Updated 3 weeks ago
- ☆17Apr 17, 2019Updated 7 years ago