Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
☆564Apr 23, 2026Updated 3 months ago
Alternatives and similar repositories for TextRL
Users that are interested in TextRL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A modular RL library to fine-tune language models to human preferences☆2,394Mar 1, 2024Updated 2 years ago
- A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)☆4,752Jan 8, 2024Updated 2 years ago
- Implementation of Reinforcement Learning from Human Feedback (RLHF)☆172Apr 7, 2023Updated 3 years ago
- A (somewhat) minimal library for finetuning language models with PPO on human feedback.☆91Nov 23, 2022Updated 3 years ago
- Train transformer language models with reinforcement learning.☆19,030Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ASR text preprocessing utility☆21Aug 5, 2024Updated 2 years ago
- A publishing website of a table collecting meta-learning-related papers in the area of human language processing.☆17Aug 2, 2021Updated 5 years ago
- 🤖📇 handling multiple nlp task in one pipeline☆57Sep 18, 2025Updated 10 months ago
- Used for adaptive human in the loop evaluation of language and embedding models.☆306Mar 1, 2023Updated 3 years ago
- ☆98May 30, 2023Updated 3 years ago
- Awesome Reinforcement Learning from Human Feedback, the secret behind ChatGPT XD☆23Dec 13, 2022Updated 3 years ago
- Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLM☆7,864Jul 27, 2026Updated last week
- official code for EMNLP21 paper☆36Dec 14, 2021Updated 4 years ago
- [NeurIPS'22 Spotlight] A Contrastive Framework for Neural Text Generation☆477Mar 7, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for the paper Fine-Tuning Language Models from Human Preferences☆1,393Jul 25, 2023Updated 3 years ago
- PFRL: a PyTorch-based deep reinforcement learning library☆1,272Mar 2, 2026Updated 5 months ago
- ☆31Jul 13, 2023Updated 3 years ago
- ☆13Sep 25, 2024Updated last year
- Code for T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5☆19Nov 29, 2022Updated 3 years ago
- 🏃 hosting nlp models in one line☆20May 8, 2024Updated 2 years ago
- Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"☆1,855Jun 17, 2025Updated last year
- ☆34Mar 25, 2023Updated 3 years ago
- Diffusion-LM☆1,245Aug 8, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for "Learning to summarize from human feedback"☆1,062Sep 5, 2023Updated 2 years ago
- Official Github repo for the paper "Evaluating the Evaluation of Diversity in Natural Language Generation"☆21Feb 23, 2021Updated 5 years ago
- one script for xls-r/xlsr/whisper fine-tuning☆42Jun 29, 2023Updated 3 years ago
- Explore different way to mix speech model(wav2vec2, hubert) and nlp model(BART,T5,GPT) together☆47Jul 3, 2025Updated last year
- A Unified Library for Parameter-Efficient and Modular Transfer Learning☆2,826Apr 26, 2026Updated 3 months ago
- A curated list of reinforcement learning with human feedback resources (continually updated)☆4,421May 20, 2026Updated 2 months ago
- ☆34Nov 17, 2021Updated 4 years ago
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Textual Style Transfer☆36Oct 2, 2022Updated 3 years ago
- [NIPS2023] RRHF & Wombat☆805Sep 22, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆26Nov 21, 2022Updated 3 years ago
- 🍳 NLPrep - dataset tool for many natural language processing task☆28Jul 30, 2021Updated 5 years ago
- Code accompanying the paper Pretraining Language Models with Human Preferences☆183Feb 13, 2024Updated 2 years ago
- This is the repo for the paper Shepherd -- A Critic for Language Model Generation☆225Aug 10, 2023Updated 2 years ago
- 聯發創新基地(MediaTek Research) 致力於研究基礎模型。我們將研究體現在適合繁體中文使用者的模型上,並在使用權許可的情況下,提供模型給學術界研究或產業界使用。☆292Apr 23, 2026Updated 3 months ago
- A Dataset for Multi-Turn Dialogue Reasoning☆334Oct 7, 2020Updated 5 years ago
- Instruction Tuning with GPT-4☆4,335Jun 11, 2023Updated 3 years ago