《多模态大模型:新一代人工智能技术范式》配套教学资源
☆328Jun 17, 2026Updated 3 months ago
Alternatives and similar repositories for Book-of-MLM
Users that are interested in Book-of-MLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [Embodied-AI-Survey-2025] Paper List and Resource Repository for Embodied AI☆2,173Jun 10, 2026Updated 3 months ago
- 大型语言模型实战指南:应用实践与场景落地☆91Sep 13, 2024Updated 2 years ago
- [IEEE T-PAMI 2023] Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering☆21Jul 6, 2023Updated 3 years ago
- 《基于BERT模型的自然语言处理实战》随书代码☆17Jun 13, 2022Updated 4 years ago
- The official repository of [CVPR2025] DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering☆30Apr 18, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [CVPR 2026] AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition☆45Apr 27, 2026Updated 5 months ago
- Embodied Question Answering (EQA) benchmark and method (ICCV 2025)☆62Aug 12, 2025Updated last year
- 🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.☆8,986Sep 21, 2026Updated last week
- Transferable Feature Representation for Visible-to-Infrared Cross-Dataset Human Action Recognition (Complexity 2018)☆13Dec 14, 2022Updated 3 years ago
- [IEEE T-CSVT 2019] Hierarchically Learned View-Invariant Representations for Cross-View Action Recognition☆15Nov 26, 2019Updated 6 years ago
- CausalVLR: A Toolbox and Benchmark for Vision-Language Causal Reasoning (多模态因果推理开源框架)☆986Oct 11, 2025Updated 11 months ago
- 通用简单工具项目☆22Oct 6, 2024Updated last year
- DDP-WM: Disentangled Dynamics Prediction for Efficient World Models (ICML-26)☆19Mar 4, 2026Updated 6 months ago
- Awesome lists of papers and codes about open-vocabulary perception, including both 3D and 2D☆65Aug 25, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [IEEE T-IP 2022] TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning☆25Dec 19, 2023Updated 2 years ago
- ☆89Jun 16, 2026Updated 3 months ago
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation (CVPR-26)☆30May 19, 2026Updated 4 months ago
- 2025.01:从零到一实现了一个多模态大模型,并命名为Reyes(睿视),R:睿,eyes:眼。Reyes的参数量为8B,视觉编码器使用的是InternViT-300M-448px-V2_5,语言模型侧使用的是Qwen2.5-7B-Instruct,Reyes也通过一个两…☆34Feb 10, 2026Updated 7 months ago
- [ICLR 2024] Official repository for "Vision-by-Language for Training-Free Compositional Image Retrieval"☆89Jul 4, 2024Updated 2 years ago
- Experimental adapter for fine-tuning Qwen3-VL as a Vision-Language-Action (VLA) model☆15Dec 21, 2025Updated 9 months ago
- The collections of MOE (Mixture Of Expert) papers, code and tools, etc.☆12Mar 15, 2024Updated 2 years ago
- The code of YOLOv5 inferencing with TensorRT C++ api is packaged into a dynamic link library , then called through Python.☆15Oct 23, 2025Updated 11 months ago
- 《从零构建AIAgent:大模型驱动的智能体设计与实战》配套代码☆26Mar 2, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆13Feb 23, 2023Updated 3 years ago
- 《大语言模型》作者:赵鑫,李军毅,周昆,唐天一,文继荣☆4,582Sep 2, 2025Updated last year
- [IEEE T-IP 2021] Semantics-aware Adaptive Knowledge Distillation for Cross-modal Action Recognition☆30Jan 6, 2025Updated last year
- The official implementation of "Cross-modal Causal Relation Alignment for Video Question Grounding. (CVPR 2025 Highlight)"☆53Apr 27, 2025Updated last year
- 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)☆25,115Jul 19, 2026Updated 2 months ago
- Implementation for What it Thinks is Important is Important: Robustness Transfers through Input Gradients (CVPR 2020 Oral)☆16Mar 24, 2023Updated 3 years ago
- Code for "Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning" (AAAI-2026 poster)☆17Mar 13, 2026Updated 6 months ago
- baseline mode for the ObjectNet competition☆18Jan 13, 2021Updated 5 years ago
- 主要记录大语言大模型(LLMs ) 算法(应用)工程师多模态相关知识☆292May 12, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Mar 1, 2022Updated 4 years ago
- 【GRSL 2021】ASEA for Change Detection☆11Sep 1, 2026Updated 3 weeks ago
- Visual Delta Generator with Large Multi-modal Model for Semi-supervised Composed Image Retrieval - CVPR2024☆21May 30, 2024Updated 2 years ago
- [CVPR 2026 Highlight 🔥] PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training☆57May 6, 2026Updated 4 months ago
- ☆14Sep 27, 2025Updated last year
- Latest Advances on Multimodal Large Language Models☆18,040Sep 18, 2026Updated last week
- Open Source Road Datasets☆19Aug 30, 2024Updated 2 years ago