A collection of the latest research and resources on Fine-Grained Multimodal Perception
☆34Jun 4, 2026Updated 3 months ago
Alternatives and similar repositories for Awesome-Fine-Grained-Multimodal-Perception
Users that are interested in Awesome-Fine-Grained-Multimodal-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] ZwZ model family: SOTA fine-grained perception performace; ZoomBench: a new challenging perception benchmark☆191May 4, 2026Updated 4 months ago
- [ICML 2026] HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling☆31May 2, 2026Updated 5 months ago
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".☆30Jul 17, 2026Updated 2 months ago
- F-16 is a powerful video large language model (LLM) that perceives high-frame-rate videos, which is developed by the Department of Electr…☆41Jul 3, 2025Updated last year
- Official implementation of "VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification"☆26May 7, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned percept…☆325Updated this week
- ArcGis二次开发大作业(C#+AE)实现一些简单的空间分析以及一些基本操作等功能☆10Dec 20, 2018Updated 7 years ago
- ☆24Apr 12, 2026Updated 5 months ago
- 科研绘图: research figure assets, posters, slides, and plotting source files.☆51Aug 2, 2026Updated 2 months ago
- Latest Advances on (RL based) Multimodal Reasoning and Generation in Multimodal LLMs☆105Sep 10, 2026Updated 3 weeks ago
- SAM4SS: Tailoring SAM and SAM2 for Semantic Segmentation☆11Jul 31, 2024Updated 2 years ago
- Real-Time ASR with CNN-BiLSTM: End-to-End Live Streaming Using PyTorch Lightning⚡☆11Jan 23, 2025Updated last year
- Code for "Neural Rendering in a Room: Amodal 3D Understanding and Free-Viewpoint Rendering for the Closed Scene Composed of Pre-Captured …☆17Oct 15, 2023Updated 2 years ago
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICLR'24] Heterogeneous Personalized Federated Learning by Local-Global Updates Mixing via Convergence Rate☆14Jun 17, 2025Updated last year
- This is the official repository of the paper "Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Schedulin…☆15Jul 27, 2025Updated last year
- ☆29Apr 8, 2025Updated last year
- [ICME 2024] DIIF (Dynamic Implicit Image Function for Efficient Arbitrary-Scale Super-Resolution).☆13Mar 13, 2024Updated 2 years ago
- 西安交通大学 beamer 主题☆20Apr 12, 2020Updated 6 years ago
- [NeurIPS 2026] Beyond SFT-to-RL: Pre-alignment via Black-BoxOn-Policy Distillation for Multimodal RL☆101May 6, 2026Updated 4 months ago
- Official dataset and pytorch implementation for the paper: `Learning Dual-Level Deformable Implicit Representation for Real-World Scale A…☆12Dec 16, 2024Updated last year
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning☆37Jun 10, 2025Updated last year
- Qt实现画图板小程序☆13Aug 19, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Jul 13, 2023Updated 3 years ago
- [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection☆30Sep 26, 2024Updated 2 years ago
- [CVPR 2026] Official release of "Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning"☆142Apr 7, 2026Updated 5 months ago
- ☆18Mar 4, 2024Updated 2 years ago
- [IEEE TMI'23] IOP-FL: Inside-Outside Personalization for Federated Medical Image Segmentation☆13Apr 2, 2023Updated 3 years ago
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]☆15Jun 1, 2026Updated 4 months ago
- Scale-aware Super-resolution Network☆19Aug 28, 2024Updated 2 years ago
- ☆10Dec 16, 2023Updated 2 years ago
- [IPMI'23] Diffusion Model based Semi-supervised Learning on Brain Hemorrhage Images for Efficient Midline Shift Quantification☆16Apr 12, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICML 2026] What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-…☆23May 15, 2026Updated 4 months ago
- [MICCAI2023] Client-Level Differential Privacy via Adaptive Intermediary in Federated Medical Imaging☆17Aug 2, 2024Updated 2 years ago
- MItosis DOmain Generalization Challenge 2021/2022☆14Jun 19, 2022Updated 4 years ago
- ☆42Updated this week
- ☆14Apr 9, 2026Updated 5 months ago
- Unofficial Implementation of Consistency Models in Pytorch☆15Mar 18, 2023Updated 3 years ago
- [ICCV 2025 Highlight] PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View☆19Jul 24, 2025Updated last year