A full pipeline to finetune Alpaca LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Alpaca architecture. Basically ChatGPT but with Alpaca
☆60Apr 28, 2023Updated 3 years ago
Alternatives and similar repositories for Alpaca-LoRA-RLHF-PyTorch
Users that are interested in Alpaca-LoRA-RLHF-PyTorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human…☆220May 20, 2024Updated 2 years ago
- Latex template for CUHK PhD Thesis☆14Jun 29, 2025Updated last year
- nlp_interview notes and answers: 该仓库主要记录 NLP 算法工程师相关的面试题和参考答案☆24Nov 16, 2023Updated 2 years ago
- LLaMA-TRL: Fine-tuning LLaMA with PPO and LoRA☆241Aug 17, 2025Updated 11 months ago
- Vietnamese GPT-J API service deployed with Docker & Helm chart☆10Dec 11, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [EMNLP 2021] PyTorch Implementation of Contrastive Domain Adaptation for Question Answering using Limited Text Corpora☆14Jul 4, 2023Updated 3 years ago
- ☆10Oct 4, 2024Updated last year
- A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.☆845Jul 1, 2024Updated 2 years ago
- ☆10Oct 31, 2022Updated 3 years ago
- Code for the SIGIR 2020 paper "A Unified Dual-view Model for Review Summarization and Sentiment Classification with Inconsistency Loss"☆21Feb 3, 2021Updated 5 years ago
- This repo contains data and code for the paper "Reasoning over Public and Private Data in Retrieval-Based Systems."☆47Jul 18, 2024Updated 2 years ago
- openai-tutorial☆14Mar 5, 2023Updated 3 years ago
- Code for ACL2024 paper - Adversarial Preference Optimization (APO).☆54Jun 3, 2024Updated 2 years ago
- Natural Language Generation by Hierarchical Decoding with Linguistic Patterns (NAACL-HLT 2018), Investigating Linguistic Pattern Ordering…☆32Sep 23, 2018Updated 7 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Finetuning LLaMA with RLHF (Reinforcement Learning with Human Feedback) based on DeepSpeed Chat☆118Jun 5, 2023Updated 3 years ago
- Bullseye Polytope Clean-Label Poisoning Attack☆18Nov 5, 2020Updated 5 years ago
- This repo is the official implementation of the ICLR'23 paper "Towards Robustness Certification Against Universal Perturbations." We calc…☆12Feb 14, 2023Updated 3 years ago
- ☆13Jul 2, 2025Updated last year
- Source code for the paper "Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data"☆20Feb 24, 2024Updated 2 years ago
- A new collection of medical VQA dataset based on MIMIC-CXR. Part of the work 'EHRXQA: A Multi-Modal Question Answering Dataset for Electr…☆101Feb 6, 2026Updated 6 months ago
- This project collects methods that enhance the comparison between AMR graphs.☆11Jun 15, 2023Updated 3 years ago
- A simple example for finetuning HuggingFace T5 model. Includes code for intermediate generation.☆26Nov 11, 2020Updated 5 years ago
- replicantlife is a framework for generative agents that can be used in a simulation engine or standalone. Agents are powered with metacog…☆34Apr 25, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆11Jul 11, 2023Updated 3 years ago
- A (somewhat) minimal library for finetuning language models with PPO on human feedback.☆91Nov 23, 2022Updated 3 years ago
- Summarization with Pointer-Generator Networks☆15Sep 1, 2020Updated 5 years ago
- Langchain_CrewAI_Gemini - An Gemini AI powered AI Agent (Multi-Agent) Project.☆14Mar 24, 2024Updated 2 years ago
- ☆12Oct 14, 2022Updated 3 years ago
- Very concise example of integrated gradients (a method to reveal areas of attention in input images)☆10Jun 17, 2019Updated 7 years ago
- Dateset Reset Policy Optimization☆30Apr 12, 2024Updated 2 years ago
- ☆17Jan 23, 2021Updated 5 years ago
- ☆17Oct 5, 2018Updated 7 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A curated list of Tsetlin Machine research☆20Nov 28, 2022Updated 3 years ago
- ☆31Sep 20, 2021Updated 4 years ago
- clip retrieval benchmark☆17May 4, 2022Updated 4 years ago
- Example Code for paper "Provably Faster Algorithms for Bilevel Optimization"☆15Dec 28, 2021Updated 4 years ago
- [NAACL 2024] Making Language Models Better Tool Learners with Execution Feedback☆43Mar 14, 2024Updated 2 years ago
- Official Implementation of "GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution"☆20Apr 3, 2024Updated 2 years ago
- Download, parse, and filter data PubMed, data-ready for The-Pile☆23Dec 16, 2021Updated 4 years ago