Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
☆564Apr 23, 2026Updated 4 months ago
Alternatives and similar repositories for TextRL
Users that are interested in TextRL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A modular RL library to fine-tune language models to human preferences☆2,394Mar 1, 2024Updated 2 years ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)☆4,755Jan 8, 2024Updated 2 years ago
- Implementation of Reinforcement Learning from Human Feedback (RLHF)☆173Apr 7, 2023Updated 3 years ago
- A (somewhat) minimal library for finetuning language models with PPO on human feedback.☆91Nov 23, 2022Updated 3 years ago
- Train transformer language models with reinforcement learning.☆19,175Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ASR text preprocessing utility☆21Aug 5, 2024Updated 2 years ago
- A publishing website of a table collecting meta-learning-related papers in the area of human language processing.☆17Aug 2, 2021Updated 5 years ago
- 🤖📇 handling multiple nlp task in one pipeline☆57Sep 18, 2025Updated 11 months ago
- Used for adaptive human in the loop evaluation of language and embedding models.☆306Mar 1, 2023Updated 3 years ago
- ☆98May 30, 2023Updated 3 years ago
- Awesome Reinforcement Learning from Human Feedback, the secret behind ChatGPT XD☆23Dec 13, 2022Updated 3 years ago
- Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLM☆7,868Jul 27, 2026Updated last month
- official code for EMNLP21 paper☆36Dec 14, 2021Updated 4 years ago
- [NeurIPS'22 Spotlight] A Contrastive Framework for Neural Text Generation☆477Mar 7, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for the paper Fine-Tuning Language Models from Human Preferences☆1,391Jul 25, 2023Updated 3 years ago
- ☆31Jul 13, 2023Updated 3 years ago
- ☆13Sep 25, 2024Updated last year
- Code for T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5☆19Nov 29, 2022Updated 3 years ago
- 🏃 hosting nlp models in one line☆20May 8, 2024Updated 2 years ago
- Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"☆1,856Jun 17, 2025Updated last year
- ☆34Mar 25, 2023Updated 3 years ago
- Diffusion-LM☆1,244Aug 8, 2024Updated 2 years ago
- Code for "Learning to summarize from human feedback"☆1,062Sep 5, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official Github repo for the paper "Evaluating the Evaluation of Diversity in Natural Language Generation"☆21Feb 23, 2021Updated 5 years ago
- one script for xls-r/xlsr/whisper fine-tuning☆42Jun 29, 2023Updated 3 years ago
- Explore different way to mix speech model(wav2vec2, hubert) and nlp model(BART,T5,GPT) together☆47Jul 3, 2025Updated last year
- A Unified Library for Parameter-Efficient and Modular Transfer Learning☆2,826Apr 26, 2026Updated 4 months ago
- A curated list of reinforcement learning with human feedback resources (continually updated)☆4,422May 20, 2026Updated 3 months ago
- ☆34Nov 17, 2021Updated 4 years ago
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Textual Style Transfer☆36Oct 2, 2022Updated 3 years ago
- [NIPS2023] RRHF & Wombat☆804Sep 22, 2023Updated 2 years ago
- ☆26Nov 21, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 🍳 NLPrep - dataset tool for many natural language processing task☆28Jul 30, 2021Updated 5 years ago
- Code accompanying the paper Pretraining Language Models with Human Preferences☆183Feb 13, 2024Updated 2 years ago
- This is the repo for the paper Shepherd -- A Critic for Language Model Generation☆225Aug 10, 2023Updated 3 years ago
- 聯發創新基地(MediaTek Research) 致力於研究基礎模型。我們將研究體現 在適合繁體中文使用者的模型上,並在使用權許可的情況下,提供模型給學術界研究或產業界使用。☆292Apr 23, 2026Updated 4 months ago
- A Dataset for Multi-Turn Dialogue Reasoning☆335Oct 7, 2020Updated 5 years ago
- Instruction Tuning with GPT-4☆4,334Jun 11, 2023Updated 3 years ago
- Convenient Text-to-Text Training for Transformers☆18Dec 10, 2021Updated 4 years ago