☆48Jul 22, 2026Updated 2 months ago
Alternatives and similar repositories for Awesome-Social-World-Models
Users that are interested in Awesome-Social-World-Models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- On Path to Multimodal Generalist: General-Level and General-Bench☆22Jul 11, 2025Updated last year
- ☆43May 15, 2025Updated last year
- 面向新同学进组的学习指南☆254Apr 26, 2026Updated 4 months ago
- This is a repository for awesome any2any work collection.☆31Jul 10, 2026Updated 2 months ago
- TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation☆70Sep 26, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Awesome papers for affective computing with llm and mllm☆39Nov 26, 2025Updated 9 months ago
- Official PyTorch Implementation for Readout Guidance, CVPR 2024☆156Jun 26, 2025Updated last year
- Official Implementation of "UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation"☆143Oct 17, 2025Updated 11 months ago
- [CVPR 2026] An official implementation of Adv-GRPO. The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image…☆90Feb 26, 2026Updated 6 months ago
- A library of long-horizon Task-and-Motion-Planning (TAMP) problems in kitchen and household scenes, as well as planners to solve them☆183May 15, 2025Updated last year
- Awesome-Emotion-Reasoning is a collection of Emotion-Reasoning works, including papers, codes and datasets☆97Dec 16, 2025Updated 9 months ago
- [CVPR 2025] Official implementation of "GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation"☆176Sep 15, 2026Updated last week
- Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation☆182Jul 17, 2025Updated last year
- [Tutorial] The Embodied Intelligence Introductory Practice of OpenMOSS Lab, SII&Fudan☆180May 27, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Leveraging Large Language Models for Visual Target Navigation☆169Oct 24, 2023Updated 2 years ago
- Visualization of DiT self attention features☆236Aug 12, 2024Updated 2 years ago
- A rebuttal editor for researchers to craft high-quality academic rebuttals — so you can focus on what to say, not how to format it.☆216Apr 3, 2026Updated 5 months ago
- MLNLP社区用来帮助大家论文Rebuttal的整理仓库。☆305Updated this week
- A Collection of Papers on Diffusion Language Models☆186Aug 2, 2026Updated last month
- This website is for the collection of VLA SOTA results.☆163Aug 31, 2026Updated 3 weeks ago
- [ICLR'25] Official code for the paper 'MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs'☆388Apr 20, 2025Updated last year
- [ICLR 2025 Oral] Seer: Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation☆310Jul 8, 2025Updated last year
- [CVPR 2025] Official code of "DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Long…☆325Mar 30, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).☆325Feb 17, 2026Updated 7 months ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆268Oct 18, 2025Updated 11 months ago
- This is the official code repo for DiT4DiT, a Vision-Action-Model (VAM) framework that combines video generation model with flow-matching…☆452Jun 30, 2026Updated 2 months ago
- [CVPRW 2026] AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation☆452Apr 13, 2025Updated last year
- ☆415Oct 20, 2025Updated 11 months ago
- Toolkits for Multimodal Emotion Recognition☆331Jun 5, 2026Updated 3 months ago
- [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of …☆506Aug 9, 2024Updated 2 years ago
- [Awesome-Spatial-VLMs] This repository is the official, community-maintained resource for the survey paper: Spatial Intelligence in Visio…☆471Jul 30, 2026Updated last month
- 🔥 OneThinker: All-in-one Reasoning Model for Image and Video [CVPR 2026]☆467Feb 28, 2026Updated 6 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- 📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.☆366Jan 8, 2026Updated 8 months ago
- 📖 从零基础到面试通关 —— 22节课彻底搞懂大语言模型 | Learn MiniMind: 系统化学习LLM训练全流程☆586Apr 1, 2026Updated 5 months ago
- Resource collection of medical agent for clinical dialogue and health☆415Feb 12, 2026Updated 7 months ago
- EMER, OV-MER (ICML25), AffectGPT (ICML25, Oral), EmoPrefer (ICLR26)☆421Feb 24, 2026Updated 6 months ago
- Training-free Regional Prompting for Diffusion Transformers 🔥☆696Nov 28, 2024Updated last year
- [RSS2024] Official implementation of "Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation"☆540Jan 19, 2026Updated 8 months ago
- [Actively Maintained🔥] A list of Embodied AI papers accepted by top conferences (ICLR, NeurIPS, ICML, RSS, CoRL, ICRA, IROS, CVPR, ICCV,…☆764May 20, 2026Updated 4 months ago