Official implementation of SAGE: a status-aware, execution-grounded planning framework that unifies temporal visual grounding, structured tool execution, and targeted error repair for egocentric interactive agents. Winner of EgoLink 2026 Track 2.
☆26Jul 21, 2026Updated this week
Alternatives and similar repositories for SAGE
Users that are interested in SAGE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This paper presents our winning submission to Subtask 2 of SemEval 2024 Task 3 on multimodal emotion cause analysis in conversations.☆25Aug 2, 2024Updated last year
- (NeXD @ CVPR 2025) Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models☆32Sep 30, 2025Updated 9 months ago
- ☆25Jul 1, 2025Updated last year
- Welcome to the official repository of Emotion-Qwen.☆27Jun 10, 2025Updated last year
- Official repository for the paper “Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models”☆28Nov 5, 2025Updated 8 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- This work introduces WordArt Designer, a user-driven framework for artistic typography synthesis, relying on Large Language Models (LLM).…☆36Jan 31, 2024Updated 2 years ago
- Code for "Improving Robustness of Vision Transformers by Reducing Sensitivity to Patch Corruptions"☆14Sep 3, 2023Updated 2 years ago
- Official Repository for VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning☆26Apr 12, 2024Updated 2 years ago
- ☆15Jun 27, 2023Updated 3 years ago
- [CVPR 2026 highlight] Official release of EgoAVU Egocentric Audio-Visual Understanding☆33Jun 8, 2026Updated last month
- [IEEE S&P'24] ODSCAN: Backdoor Scanning for Object Detection Models☆22Oct 5, 2025Updated 9 months ago
- Unofficial implementation of "Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition - CVPR'21"☆19Mar 3, 2024Updated 2 years ago
- A Python project that parses Markdown files into a tree structure, then processes them into semantically meaningful text chunks.☆22Jan 9, 2025Updated last year
- Official Code of "Imperceptible Adversarial Attack via Invertible Neural Networks"☆24Jul 24, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- SMILE: A Multimodal Dataset for Understanding Laughter☆13Jun 15, 2023Updated 3 years ago
- Official implementation for "Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling"☆10May 21, 2024Updated 2 years ago
- Code corresponding to the paper: "On the Robustness of Vision Transformers": https://arxiv.org/abs/2104.02610☆25Dec 16, 2025Updated 7 months ago
- [ICLR 2026] VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models☆21Feb 22, 2026Updated 5 months ago
- Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models (ACL-Findings 2024)☆16Apr 23, 2024Updated 2 years ago
- TheraMind : A Strategic and Adaptive Agent for Longitudinal Psychological Counseling (WWW 2026)☆25May 15, 2026Updated 2 months ago
- This is the official implementation of our paper Untargeted Backdoor Attack against Object Detection.☆27Mar 6, 2023Updated 3 years ago
- My slides and examples for bachelor deep learning course☆12Jun 2, 2022Updated 4 years ago
- Code for "Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations" (CVPR 2024 Oral)☆19Jun 23, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for near-shore thermal water segmentation on UAVs. Datasets for network training included.☆17Dec 18, 2023Updated 2 years ago
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning (ACM MM 2025 Oral)☆35Oct 11, 2025Updated 9 months ago
- Multiplatform Command Line Interface for OneDrive☆14Mar 30, 2025Updated last year
- SAM-FNet: SAM-Guided Fusion Network for Laryngo-Pharyngeal Tumor Detection☆12Aug 10, 2024Updated last year
- ☆20Mar 12, 2025Updated last year
- Give us minutes, we give back a faster Mamba. The official implementation of "Faster Vision Mamba is Rebuilt in Minutes via Merged Token …☆39Dec 18, 2024Updated last year
- Decoupled Kullback-Leibler Divergence Loss (DKL), NeurIPS 2024 / Generalized Kullback-Leibler Divergence Loss (GKL), TPAMI 2026☆51Jun 17, 2026Updated last month
- ☆131Jun 11, 2026Updated last month
- A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor☆48Jan 13, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [NeurIPS 2023 Spotlight] In-Context Impersonation Reveals Large Language Models' Strengths and Biases☆22Nov 30, 2024Updated last year
- ☆22Jan 14, 2023Updated 3 years ago
- [NeurIPS2020] The official repository of "AdvFlow: Inconspicuous Black-box Adversarial Attacks using Normalizing Flows".☆49Oct 3, 2023Updated 2 years ago
- Applying Octave UNet for Retinal Vessel Segmentation.☆22Dec 25, 2020Updated 5 years ago
- RLLaVA is a user-friendly framework for multi-modal RL research and optimized for resource-constrained teams.☆58Mar 18, 2026Updated 4 months ago
- a training-free approach to accelerate ViTs and VLMs by pruning redundant tokens based on similarity☆44May 24, 2025Updated last year
- we propose FlexEdit, an end-to-end image editing method that leverages both free-shape masks and language instructions for Flexible Editi…☆32Aug 22, 2024Updated last year