The official repository for CVPR'26 Paper "APPO: Attention-guided Perception Policy Optimization for Video Reasoning"
☆16Mar 19, 2026Updated 5 months ago
Alternatives and similar repositories for APPO
Users that are interested in APPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆17Oct 25, 2024Updated last year
- ☆36Jul 9, 2025Updated last year
- [CVPR 2025] Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation☆86Dec 24, 2025Updated 8 months ago
- A python implement for Certifiable Robust Multi-modal Training☆21Jun 21, 2025Updated last year
- Spatial-Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation (TIP 2024, ACM MM 2023)☆19Mar 13, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The official code of Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval (AAAI2024)☆32Mar 29, 2024Updated 2 years ago
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆30Apr 17, 2026Updated 4 months ago
- PatchBackdoor is a code base associated with paper PatchBackdoor.☆12Aug 27, 2024Updated 2 years ago
- [CVPR 2026] VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking☆69Mar 23, 2026Updated 5 months ago
- How Robust are Randomized Smoothing based Defenses to Data Poisoning? (CVPR 2021)☆14Jul 16, 2021Updated 5 years ago
- [ACL 2025] PruneVid: Visual Token Pruning for Efficient Video Large Language Models☆71May 15, 2025Updated last year
- Automatically update arXiv papers about SOT & VLT, Multi-modal Learning, LLM and Video Understanding using Github Actions.☆48Sep 7, 2026Updated last week
- The repo for "Enhancing Multi-modal Cooperation via Sample-level Modality Valuation", CVPR 2024☆62Nov 5, 2024Updated last year
- [ICML 2026] Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning☆135Jul 2, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official code repository of Shuffle-R1☆26Feb 23, 2026Updated 6 months ago
- EVA: Efficient Reinforcement Learning for End-to-End Video Agent☆26May 6, 2026Updated 4 months ago
- Official implementation of paper "OED: Towards One-stage End-to-End Dynamic Scene Graph Generation".☆31Mar 26, 2024Updated 2 years ago
- ☆20Jun 26, 2026Updated 2 months ago
- ☆13Jul 15, 2024Updated 2 years ago
- [CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice☆89Feb 27, 2026Updated 6 months ago
- An environment for testing tensorflow models.☆14Oct 29, 2017Updated 8 years ago
- The offical repo for "Play to the Score: Stage-Guided Dynamic Multi-Sensory Fusion for Robotic Manipulation", CoRL 2024 (ORAL)☆22Jun 25, 2025Updated last year
- [ICLR 2025] IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model☆37Nov 27, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆22Jul 25, 2024Updated 2 years ago
- ☆18Feb 8, 2026Updated 7 months ago
- Multi-Granularity Language-Guided Multi-Object Tracking☆26Nov 3, 2025Updated 10 months ago
- [NeurIPS'24] MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts☆19Oct 7, 2024Updated last year
- All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment☆21Feb 11, 2025Updated last year
- ☆12Mar 22, 2025Updated last year
- ☆56May 16, 2026Updated 3 months ago
- ☆25Dec 23, 2024Updated last year
- Pytorch official implementation for Imitating Unknown Policies via Exploration.☆14Oct 3, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICCV 2025] Dynamic-VLM☆28Dec 16, 2024Updated last year
- Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning☆46Mar 2, 2026Updated 6 months ago
- ☆14Mar 29, 2023Updated 3 years ago
- Official implementation for "Mixture of In-Context Experts Enhance LLMs’ Awareness of Long Contexts" (Accepted by Neurips2024)☆14Jan 7, 2025Updated last year
- ☆15Feb 28, 2023Updated 3 years ago
- Model LEGO: Creating Models Like Disassembling and Assembling Building Blocks☆17Jan 15, 2025Updated last year
- Temporal Sentence Grounding in Videos / Natural Language Video Localization / Video Moment Retrieval的相关工作☆31Mar 4, 2022Updated 4 years ago