A full pipeline to finetune ChatGLM LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the ChatGLM architecture. Basically ChatGPT but with ChatGLM
☆138Apr 28, 2023Updated 3 years ago
Alternatives and similar repositories for ChatGLM-LoRA-RLHF-PyTorch
Users that are interested in ChatGLM-LoRA-RLHF-PyTorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 对ChatGLM直接使用RLHF提升或降低目标输出概率|Modify ChatGLM output with only RLHF☆195May 23, 2023Updated 3 years ago
- A full pipeline to finetune Alpaca LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human…☆60Apr 28, 2023Updated 3 years ago
- NCA-SEM module for Jamovi. Necessary Condition Analysis via Structural Equation Modeling (NCA-SEM) is a data analysis method that is used…☆15Jun 24, 2026Updated 2 months ago
- A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human…☆220May 20, 2024Updated 2 years ago
- ☆43Dec 15, 2023Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- 微调ChatGLM☆128May 5, 2023Updated 3 years ago
- moss chat finetuning☆51Apr 23, 2024Updated 2 years ago
- LLaMA-TRL: Fine-tuning LLaMA with PPO and LoRA☆238Aug 17, 2025Updated last year
- ZYN: Zero-Shot Reward Models with Yes-No Questions☆34Aug 15, 2023Updated 3 years ago
- ChatGLM-6B添加了RLHF的实现,以及部分核心代码的逐行讲解 ,实例部分是做了个新闻短标题的生成,以及指定context推荐的RLHF的实现☆88Aug 16, 2023Updated 3 years ago
- chatglm-6b微调/LORA/PPO/推理, 样本为自动生成的整数/小数加减乘除运算, 可gpu/cpu☆165Aug 24, 2023Updated 3 years ago
- nlp_interview notes and answers: 该仓库主要记录 NLP 算法工程师相关的面试题和参考答案☆24Nov 16, 2023Updated 2 years ago
- Finetuning LLaMA with RLHF (Reinforcement Learning with Human Feedback) based on DeepSpeed Chat☆117Jun 5, 2023Updated 3 years ago
- Humanable Chat Generative-model Fine-tuning | LLM微调☆205Sep 22, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ChatGLM2-6B 全参数微调,支持多轮对话的高效微调。☆400Aug 17, 2023Updated 3 years ago
- Secrets of RLHF in Large Language Models Part I: PPO☆1,430Mar 3, 2024Updated 2 years ago
- this is an implementation for the paper Improve Mathematical Reasoning in Language Models by Automated Process Supervision from google de…☆50Jul 8, 2025Updated last year
- ⭐️ NLP Algorithms with transformers lib. Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SF…☆2,423Sep 29, 2023Updated 2 years ago
- Training a reward model for RLHF using RWKV.☆15Jun 5, 2023Updated 3 years ago
- 基于ChatGLM-6B + LoRA的Fintune方案☆3,738Nov 25, 2023Updated 2 years ago
- PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation☆16Mar 28, 2023Updated 3 years ago
- MarketGPT: Developing a Pre-trained transformer (GPT) for Modeling Financial Time Series☆19Sep 5, 2025Updated last year
- A Toolkit for Fine-Tuning Large Language Models with LoRA and DeepSpeed☆11Apr 14, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 用RLHF可选LoRA对LLaMA和MOSS进行训练|Training LLaMA or MOSS with RLHF [LoRA]☆21May 16, 2023Updated 3 years ago
- chatglm 6b finetuning and alpaca finetuning☆1,521Mar 9, 2025Updated last year
- ChatGLM2-6B微调, SFT/LoRA, instruction finetune☆106Jul 19, 2023Updated 3 years ago
- Instruction Tuning with GPT-4☆4,331Jun 11, 2023Updated 3 years ago
- ☆23Jun 23, 2023Updated 3 years ago
- 基于ChatGLM-6B、ChatGLM2-6B、ChatGLM3-6B模型,进行下游具体任务微调,涉及Freeze、Lora、P-tuning、全参微调等☆2,771Dec 12, 2023Updated 2 years ago
- Collection of links, tutorials and best practices of how to collect the data and build end-to-end RLHF system to finetune Generative AI m…☆227Jul 24, 2023Updated 3 years ago
- ChatGLM-6B 指令学习|指令数据|Instruct☆650Apr 10, 2023Updated 3 years ago
- Implementation of Chinese ChatGPT☆285Nov 20, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 实现Blip2RWKV+QFormer的多模态图文对话大模型,使用Two-Step Cognitive Psychology Prompt方法,仅3B参数的模型便能够出现类人因果思维链。对标MiniGPT-4,ImageBind等图文对话大语言模型,力求以更小的算力和资源实…☆41Jul 17, 2023Updated 3 years ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)☆4,754Jan 8, 2024Updated 2 years ago
- LLaMa Tuning with Stanford Alpaca Dataset using Deepspeed and Transformers☆50Mar 15, 2023Updated 3 years ago
- This is the oficial repository for "Safer-Instruct: Aligning Language Models with Automated Preference Data"☆17Feb 22, 2024Updated 2 years ago
- A curated list of reinforcement learning with human feedback resources (continually updated)☆4,426May 20, 2026Updated 4 months ago
- Python scripts for setting up private LLM's on local and in the cloud with LangChain, GPT4All and Cerebrium☆12May 29, 2023Updated 3 years ago
- Official PyTorch implementation of TTS Style Transfer☆25Jun 22, 2022Updated 4 years ago