A travel agent based on Qwen2.5, fine-tuned by SFT + DPO/PPO/GRPO using traveling question-answer dataset, a mindmap can be output using the response. A RAG system is build upon the tuned qwen2, using Prompt-Template + Tool-Use + Chroma embedding database + LangChain
☆83Jul 6, 2026Updated 2 months ago
Alternatives and similar repositories for Travel-Agent-based-on-Qwen2-RLHF
Users that are interested in Travel-Agent-based-on-Qwen2-RLHF are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 基于Qwen2+SFT+DPO的医疗问答系统,项目中使用了自定义的 SFTTrainer/DPOTrainer/TRPOTrainer用于训 练,其次,项目还调用各种知识库工具(neo4j, milvus, LDA, 等)进行自动化训练数据生成。另外,使用 vllm 用于推理…☆88Apr 29, 2026Updated 4 months ago
- 超简单使用监督微调SFT和强化学习RL去训练领域Agent☆35Oct 20, 2025Updated 10 months ago
- [CVPR'2025] Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions☆19Jan 16, 2026Updated 7 months ago
- AFAC2024金融智能创新大赛☆68Nov 27, 2024Updated last year
- CCKS2023-PromptCBLUE: Code implement of TianChi completition☆21Feb 27, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A comparison of deepseek grpo and qwen gspo on Qwen2.5-1.5B-Instruct fine tunning.☆170Mar 28, 2026Updated 5 months ago
- 为加快推动人工智能及大模型技术在金融科技领域的应用转化,助力挖掘优质创新创业方案,强化科技人才高地聚集氛围,在上海市科学技术委员会指导下,国内外20多家顶尖高校、学术机构和金融科技企业联合发起“AFAC2025金融智能创新大赛”。☆39Nov 4, 2025Updated 10 months ago
- ☆18Mar 15, 2026Updated 5 months ago
- ☆12Sep 4, 2026Updated last week
- Qwen2-VL在文旅领域的LLaMA-Factory微调案例 The case for fine-tuning Qwen2-VL in the field of historical literature and museums☆15Sep 17, 2024Updated last year
- Finance specialized RAG System for the ACM-ICAIF '24 Competition.☆61Nov 27, 2024Updated last year
- 一个低成本、易于上手的多模态大模型学习项目。基于Qwen3-0.6B和CLIP构建,使用LLaVA架构和LoRA微调,在消费级16G显卡上数小时即可完成训练☆55Sep 15, 2025Updated 11 months ago
- PipelineLLM 是一个系统性的大语言模型(LLM)后训练学习项目,涵盖从监督微调(SFT)到偏好优化(DPO)、强化学习(RLHF/PPO/GRPO)再到持续学习(Continual Learning)的完整技术栈。☆35Jan 16, 2026Updated 7 months ago
- ☆10Apr 7, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 一个包含了多种主流大模型微调 方案的实战代码库,基于Qwen3系列模型☆141Aug 10, 2025Updated last year
- 用大模型批量处理数据,现支持各种大模型做OCR,支持通义千问, 月之暗面, 百度飞桨OCR, OpenAI 和LLAVA。Use LLM to generate or clean data for academic use. Support OCR with qwen, m…☆18Sep 15, 2024Updated last year
- This project involves code for summarization of legal documents with a two step process. First we fine-tune the llm for legal documents u…☆21Jan 26, 2025Updated last year
- 此仓库是我们小组在《计算机游戏开发》课程(深圳大学)的大作业,是一个模仿《slay the spire》的卡牌游戏☆10Jun 28, 2019Updated 7 years ago
- 基于PaddleNLP的对话意图识别☆10Apr 11, 2023Updated 3 years ago
- 拼好RAG:手搓并融合了GraphRAG、LightRAG、Neo4j-llm-graph-builder进行知识图谱构建以及搜索;整合DeepSearch技术实现私域RAG的推理;自制针对GraphRAG的评估框架| Integrate GraphRAG, LightRA…☆2,331Nov 5, 2025Updated 10 months ago
- ☆19May 17, 2025Updated last year
- ☆14May 23, 2025Updated last year
- Unify Efficient Fine-tuning of RAG Retrieval, including Embedding, ColBERT, ReRanker.☆23Aug 28, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆16Oct 3, 2022Updated 3 years ago
- [CVPR 2025] MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities☆31Apr 6, 2025Updated last year
- 本项目对Deepseek-R1-Distill-Qwen-7B进行心理咨询CoT数据的LoRA微调,以进一步提升Deepseek-R1-Distill-Qwen-7B在心理咨询领域的慢思考能力。☆12Mar 11, 2025Updated last year
- ParetoDrug☆12Sep 3, 2024Updated 2 years ago
- PyTorch implementation of the paper "NestE: Modeling Nested Relational Structures for Knowledge Graph Reasoning" (AAAI'24)☆14Jul 5, 2024Updated 2 years ago
- ☆13Dec 29, 2021Updated 4 years ago
- 本项目由三个模块构成。意图识别:判断用户的意图是业务型还是闲聊型;模型检索:该部分构建一个语料库,当用户 发起新的query(通过意图识别判断为业务型对话)时,为用户匹配query检索的最佳response,使用HSWN进行召回(粗排), 然后构建句子的相似度,并利用Lig…☆12Feb 18, 2021Updated 5 years ago
- Multivariate Time Series Anomaly Detection by Capturing Coarse-Grained Intra- and Inter-Variate Dependencies☆17Feb 13, 2025Updated last year
- ☆25Oct 22, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- AiMed面向中文医学的人工智能大语言模型期望实现有效处理医学知识问答、医学论文阅读、医学文献检索等任务和在医学科研中的应用。☆13Feb 8, 2025Updated last year
- Qwen3 Fine-tuning: Medical R1 Style Chat☆334May 31, 2025Updated last year
- Demo app with Loguru logging, async middleware to generate X-request-Id. Works with Gunicorn or Uvicorn, and is safe to use with async/th…☆10Feb 2, 2022Updated 4 years ago
- 2024CCF国际AIOps挑战赛-赛道二(GLM4):基于检索增强的运维知识问答挑战赛解决方案分享。☆14Jul 5, 2024Updated 2 years ago
- [ACM TOIS] Multi-Behavior Recommendation with Personalized Directed Acyclic Behavior Graphs☆13Dec 6, 2024Updated last year
- ACM Transactions on Information Systems (TOIS), the code and datasets for CKML.☆13Aug 31, 2023Updated 3 years ago
- ☆23Mar 20, 2024Updated 2 years ago