从零预训练LLM、SFT、RLHF、DPO笔记整理+面试问题
☆21Sep 2, 2024Updated 2 years ago
Alternatives and similar repositories for LLM_Learning_ph
Users that are interested in LLM_Learning_ph are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [原理解析] 大模型基本功(手撕Transformer模型、手撕PPO、GRPO、DPO训练器)☆33Jul 8, 2025Updated last year
- The implementations of paper "Reinforced Preference Optimization for Recommendation" (ReRe).☆23Sep 10, 2026Updated last month
- 🔥🔥🔥 基于 PyTorch Lightning 和 MIND 数据集的模块化新闻推荐系统框架。实现了从特征工程到召回 (DSSM) 与排序 (Deep, DCN, WideDeep, FM) 的完整链路。☆49Apr 12, 2026Updated 5 months ago
- 本项目是一个基于LangChain构建的多Agent系统,结合Streamlit实现的Web界面,能够根据用户输入进行网络搜索并提供旅游相关的聊天服务。此外,该系统还具备基于本地知识库的推销功能,为用户提供个性化的旅游产品推荐。☆16Apr 20, 2025Updated last year
- C++ implementation of FRAIGs. Won the 1st place in 2018 Cadence-sponsored contest in NTU DSnP.☆10Oct 21, 2020Updated 5 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆13Aug 9, 2023Updated 3 years ago
- ☆12Jun 23, 2023Updated 3 years ago
- 基于Qwen2+SFT+DPO的医疗问答系统,项目中使用了自定义的 SFTTrainer/DPOTrainer/TRPOTrainer用于训练,其次,项目还调用各种知识库工具(neo4j, milvus, LDA, 等)进行自动化训练数据生成。另外,使用 vllm 用于推理…☆90Apr 29, 2026Updated 5 months ago
- 基于 Qwen/Qwen3-0.6B 的医疗问答微调与推理示例项目,记录健康数据、生成建议与报告☆16Sep 30, 2025Updated last year
- 开源AGENT代码分享☆16Jul 26, 2025Updated last year
- Sequence Tagging for Biomedical Extractive Question Answering (Bioinformatics'2020)☆10Jul 3, 2023Updated 3 years ago
- 这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。☆63Apr 13, 2025Updated last year
- Code for the 2025 ACL publication "Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs"☆33Jun 25, 2025Updated last year
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 基于BERT-MRC(阅读理解)的命名实体识别模型☆20Mar 15, 2022Updated 4 years ago
- 基于DeepSpeed的大模型微调教程,详细介绍如何使用DeepSpeed进行微调和分布式训练文本总结大模型☆19May 6, 2026Updated 5 months ago
- LangGraph agent template with MCP.☆29Apr 8, 2025Updated last year
- LLM手撕代码合集☆24Mar 25, 2025Updated last year
- PipelineLLM 是一个系统性的大语言模型(LLM)后训练学习项目,涵盖从监督微调(SFT)到偏好优化(DPO)、强化学习(RLHF/PPO/GRPO)再到持续学习(Continual Learning)的完整技术栈。☆37Jan 16, 2026Updated 8 months ago
- Code for Analyzing Redundancy in Pretrained Transformer Models accepted at EMNLP 2020☆14Oct 6, 2020Updated 6 years ago
- 阿里天池智慧交通预测挑战赛-Top7 /1716队☆14Jun 9, 2018Updated 8 years ago
- CoCoFL: Communication- and Computation-Aware Federated Learning via Partial NN Freezing and Quantization☆12Aug 3, 2024Updated 2 years ago
- 基于 OneKE 的知识图谱构建与 RAG 问答系统搭建☆27Jun 29, 2024Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- This repo contains some extensions of deepspeed-chat for fine-tuning LLMs (SFT+RLHF).☆21Jul 2, 2024Updated 2 years ago
- DTVis:交通流量时空演变特征可视分析;数据源:2019CCF BDCI-可视化大赛☆17Jul 15, 2021Updated 5 years ago
- Code and data of WSDM 2023 paper "Hansel: A Chinese Few-Shot and Zero-Shot Entity Linking Benchmark".☆24Jun 1, 2023Updated 3 years ago
- 简单实现了一下基于知识图谱和文本文档联合做检索增强(RAG)大模型的实现,这里采用的数据分别是管廊维护领域的文本文档和专家知识图谱☆24Jun 6, 2024Updated 2 years ago
- 本项目将基于多模态,RAG以及LLM等技术,打造了一个基于手相算命的系统☆30Aug 28, 2024Updated 2 years ago
- A very simple navigational search homepage with a background using Bing's image API and support for adding search engines on your own.☆10Jan 19, 2026Updated 8 months ago
- Bert + PCNN and PCNN 中文关系抽取任务☆20Dec 30, 2022Updated 3 years ago
- 2018年研究生室友推荐系统——Roommate Matching——简单小应用帮助同学寻找习性相同的室友☆11Apr 3, 2019Updated 7 years ago
- Simple implementation of Retrieval-Augmented Generation System☆28Oct 24, 2024Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [CVPR2023] Practical Network Acceleration with Tiny Sets☆13Jul 28, 2023Updated 3 years ago
- 智能客服系统架构与多Agent协作☆34Oct 9, 2025Updated last year
- Experiments codes for WSDM '24 paper "MultiFS: Automated Multi-Scenario Feature Selection in Deep Recommender Systems"☆11May 31, 2024Updated 2 years ago
- [NeurIPS2022] Where to Pay Attention in Sparse Training for Feature Selection?☆13Feb 10, 2023Updated 3 years ago
- 2021 语言与智能技术竞赛关系 篇章级关系抽取☆17Sep 8, 2021Updated 5 years ago
- Q-RR, DIANA-RR, Q-NASTYA, NASTYA-DIANA, QSGD, DIANA, FedCOM and FedPAQ on logistic loss with L2 regularization☆11Nov 1, 2022Updated 3 years ago
- [Findings of EMNLP22] From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models☆19Mar 16, 2023Updated 3 years ago