Train a model for Image Caption from ViT and GPT pretrained model
☆18Mar 25, 2023Updated 3 years ago
Alternatives and similar repositories for Chinese-Image-Caption
Users that are interested in Chinese-Image-Caption are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Mar 8, 2025Updated last year
- 基于CLIP实现以文精准搜图☆16Sep 20, 2023Updated 2 years ago
- Implementation of our AAAI2022 paper, Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text Matching.☆36Jun 16, 2023Updated 3 years ago
- KERL: Knowledge-Enhanced Personalized Recipe Recommendation using Large Language Models☆15Jun 18, 2025Updated last year
- A reimplementation of KOSMOS-1 from "Language Is Not All You Need: Aligning Perception with Language Models"☆27Mar 3, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Learnable Pillar-based Re-ranking for Image-Text Retrieval. SIGIR'23☆22Jul 31, 2023Updated 2 years ago
- 吃啥好呢 - 个性化美食推荐☆19Apr 13, 2026Updated 3 months ago
- mysql node.js 基于协同过滤美食推荐系统-数据库☆11Feb 19, 2019Updated 7 years ago
- ☆17Sep 23, 2024Updated last year
- record and share my reading everyday☆12Apr 1, 2016Updated 10 years ago
- Efficient Token-Guided Image-Text Retrieval with Consistent Multimodal Contrastive Training☆30Jun 20, 2023Updated 3 years ago
- This repository contains the code for our paper: Enhancing Abnormality Grounding for Vision-Language Models with Knowledge Descriptions☆19Jun 24, 2025Updated last year
- A web application that recommends songs via "country arithmetic" and hand-rolled Implicit Matrix Factorization☆10May 5, 2017Updated 9 years ago
- This repository contains the code accompanying the paper "A Self-Guided Framework for Radiology Report Generation", accepted by MICCAI 20…☆20Mar 11, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Image Captioning using combination of object detection via YOLOv5 and Encoder Decoder LSTM model☆15Oct 13, 2022Updated 3 years ago
- 基于YOLOv11 + FastAPI + Vue + Docker的目标检测可视化网页☆17Mar 3, 2026Updated 4 months ago
- Joint Embedding of Deep Visual and Semantic Features for Medical Image Report Generation☆18Nov 13, 2025Updated 8 months ago
- 句子匹配模型,包括无监督的SimCSE、ESimCSE、PromptBERT,和有监督的SBERT、CoSENT。☆98Oct 29, 2022Updated 3 years ago
- ⛏️This is the storage of my Slides、Reports and Papers. | 存储PPT、报告和论文☆12Oct 27, 2024Updated last year
- 小鹿健康是一个面向高校学生与年轻群体的健康管理App,集健康数据记录、饮食管理、AI分析推荐与社区互动等于一体。配合大模型智能体、图像识别与健康知识图谱,致力于打造一个真正智能化、个性化的健康生活APP。☆27Jun 27, 2025Updated last year
- ☆14May 5, 2019Updated 7 years ago
- Retrieval Augmented Generation demo using Microsoft's phi-2 LLM and langchain☆19Feb 12, 2024Updated 2 years ago
- ☆11Nov 21, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A 3D Slicer extension for FastSAM3D☆22Aug 31, 2024Updated last year
- 暑期研究:神经网络解偏微分方程(Neural networks for solving differential equations)☆14Apr 27, 2019Updated 7 years ago
- ☆11Sep 18, 2020Updated 5 years ago
- 看图说话机器人☆30Mar 18, 2019Updated 7 years ago
- Fetching confused chars, including same pronunciation, similar pronunciation and similar character pattern☆21Jan 20, 2023Updated 3 years ago
- ☆66Dec 15, 2023Updated 2 years ago
- 回声Echo:AI文案助手☆10May 6, 2023Updated 3 years ago
- A new novel multi-modality (Vision) RAG architecture☆39Oct 1, 2024Updated last year
- Multi-span Style Extraction for Generative Reading Comprehension☆10Apr 2, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An LLM answering questions about your IFC file by querying a graph database that represents the file.☆27Apr 23, 2025Updated last year
- ChatTTS is a generative speech model for daily dialogue.☆14Oct 21, 2024Updated last year
- Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal Hashing☆37Dec 9, 2020Updated 5 years ago
- This repo offers advanced tutorials for LLMs, BERT-based models, and multimodal models, covering fine-tuning, quantization, vocabulary ex…☆24May 5, 2025Updated last year
- Code and data for ACM MM '23 paper “MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark Evaluation”☆30Aug 20, 2024Updated last year
- Enhance your knowledge in medical research with the help of LLM and RAG.☆35Oct 24, 2024Updated last year
- ☆14Nov 12, 2021Updated 4 years ago