Official implementation of SAGE: a status-aware, execution-grounded planning framework that unifies temporal visual grounding, structured tool execution, and targeted error repair for egocentric interactive agents. Winner of EgoLink 2026 Track 2.
☆46Aug 16, 2026Updated last month
Alternatives and similar repositories for SAGE
Users that are interested in SAGE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repository of Orchestra-o1: Omnimodal Agent Orchestration☆170Jun 15, 2026Updated 3 months ago
- This paper presents our winning submission to Subtask 2 of SemEval 2024 Task 3 on multimodal emotion cause analysis in conversations.☆24Aug 2, 2024Updated 2 years ago
- ☆24Jul 1, 2025Updated last year
- Welcome to the official repository of Emotion-Qwen.☆26Aug 12, 2026Updated last month
- Official repository for the paper “Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models”☆30Nov 5, 2025Updated 10 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning☆617Aug 28, 2026Updated 3 weeks ago
- From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health (TAFFC)☆125Aug 5, 2026Updated last month
- Beyond Empathy: Integrating Diagnostic and Therapeutic Reasoning with Large Language Models for Mental Health Counseling☆54Apr 19, 2026Updated 5 months ago
- Code for "Improving Robustness of Vision Transformers by Reducing Sensitivity to Patch Corruptions"☆14Sep 3, 2023Updated 3 years ago
- 可逆水印论文复现☆10Nov 30, 2019Updated 6 years ago
- EmoCapCLIP: Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions☆22Jul 29, 2025Updated last year
- Official Repository for VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning☆26Apr 12, 2024Updated 2 years ago
- ☆31Jul 24, 2026Updated last month
- [CVPR 2026 highlight] Official release of EgoAVU Egocentric Audio-Visual Understanding☆35Jun 8, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [IEEE S&P'24] ODSCAN: Backdoor Scanning for Object Detection Models☆22Oct 5, 2025Updated 11 months ago
- ☆28Apr 29, 2025Updated last year
- "FORB: A Flat Object Retrieval Benchmark for Universal Image Embedding", NeurIPS 2023 Datasets and Benchmarks Track☆13Jun 20, 2024Updated 2 years ago
- SMILE: A Multimodal Dataset for Understanding Laughter☆13Jun 15, 2023Updated 3 years ago
- Official implementation for "Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling"☆10May 21, 2024Updated 2 years ago
- Code corresponding to the paper: "On the Robustness of Vision Transformers": https://arxiv.org/abs/2104.02610☆25Dec 16, 2025Updated 9 months ago
- [Neurips 2024] Disentangled Graph Homophily☆29Jan 21, 2025Updated last year
- A Unimodal Valence-Arousal Driven Contrastive Learning Framework for Multimodal Multi-Label Emotion Recognition (ACM MM 2024 oral)☆31Nov 4, 2024Updated last year
- 在手写数字集MNIST上使用变分自动编码器作为encoder和decoder的ldm☆24May 23, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TheraMind : A Strategic and Adaptive Agent for Longitudinal Psychological Counseling (WWW 2026)☆28May 15, 2026Updated 4 months ago
- My slides and examples for bachelor deep learning course☆12Jun 2, 2022Updated 4 years ago
- Code for "Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations" (CVPR 2024 Oral)☆20Jun 23, 2024Updated 2 years ago
- Code for near-shore thermal water segmentation on UAVs. Datasets for network training included.☆17Dec 18, 2023Updated 2 years ago
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning (ACM MM 2025 Oral)☆42Oct 11, 2025Updated 11 months ago
- ☆11Mar 3, 2024Updated 2 years ago
- [ICLR 2025] Let Your Features Tell The Differences: Understanding Graph Convolution By Feature Splitting☆16Nov 24, 2025Updated 9 months ago
- A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor☆58Jan 13, 2026Updated 8 months ago
- [NeurIPS 2023 Spotlight] In-Context Impersonation Reveals Large Language Models' Strengths and Biases☆22Nov 30, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for "DAMEX: Dataset-aware Mixture-of-Experts for visual understanding of mixture-of-datasets", accepted at Neurips 2023 (Main confer…☆28Mar 29, 2024Updated 2 years ago
- RLLaVA is a user-friendly framework for multi-modal RL research and optimized for resource-constrained teams.☆57Updated this week
- we propose FlexEdit, an end-to-end image editing method that leverages both free-shape masks and language instructions for Flexible Editi…☆32Aug 22, 2024Updated 2 years ago
- This repositary contains an implemetation of the two stage networks CVNet and SuperGlobal, for Image Retrieval.☆24Feb 20, 2024Updated 2 years ago
- Official Code for the WWW'24 Paper: "Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models"☆26Apr 16, 2025Updated last year
- [ECCV 2024] Official implementation of the paper "Towards Latent Masked Image Modeling for Self-Supervised Visual Representation Learning…☆31Mar 5, 2025Updated last year
- Codebase for the work “Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?”☆77Apr 14, 2026Updated 5 months ago