Official repository for "BLEUBERI: BLEU is a surprisingly effective reward for instruction following"
☆32Jun 5, 2025Updated last year
Alternatives and similar repositories for BLEUBERI
Users that are interested in BLEUBERI are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆62Sep 24, 2024Updated last year
- Vocabulary Parallelism☆26Mar 10, 2025Updated last year
- IntructIR, a novel benchmark specifically designed to evaluate the instruction following ability in information retrieval models. Our foc…☆32Jun 13, 2024Updated 2 years ago
- ☆10Nov 8, 2023Updated 2 years ago
- ☆12Jun 5, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Consistent dialogue generation☆16Oct 26, 2022Updated 3 years ago
- Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals☆11Jan 8, 2026Updated 7 months ago
- Implementation for the paper "Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning"☆11Jan 10, 2025Updated last year
- An Empirical Study On Contrastive Search And Contrastive Decoding For Open-ended Text Generation☆27Jun 7, 2024Updated 2 years ago
- Official repo for Learning to Reason for Long-Form Story Generation☆78Apr 19, 2025Updated last year
- [ICML2025] Official Repo for Paper "Optimizing Temperature for Language Models with Multi-Sample Inference"☆23Feb 16, 2025Updated last year
- Danmuku dataset☆12Jul 7, 2023Updated 3 years ago
- [ICML 2025] Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment (https://arxiv.org/abs/2410.02197)☆44Jun 15, 2026Updated 2 months ago
- Code for "Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations" [NAACL Findings 2024]☆14Apr 3, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ECCV 2024] Official PyTorch implementation of LUT "Learning with Unmasked Tokens Drives Stronger Vision Learners"☆15Dec 1, 2024Updated last year
- An unofficial implementation of SOLAR-10.7B model and the newly proposed interlocked-DUS(iDUS) implementation and experiment details.☆14Mar 20, 2024Updated 2 years ago
- This repository contains the code used for the experiments in the paper "Language Models use Lookbacks to Track Beliefs".☆17Mar 14, 2026Updated 5 months ago
- ☆16Jul 23, 2024Updated 2 years ago
- Code for the paper-"Mirostat: A Perplexity-Controlled Neural Text Decoding Algorithm" (https://arxiv.org/abs/2007.14966).☆62Feb 7, 2022Updated 4 years ago
- Self-Alignment with Principle-Following Reward Models☆170Sep 18, 2025Updated 11 months ago
- Learning from preferences is a common paradigm for fine-tuning language models. Yet, many algorithmic design decisions come into play. Ou…☆32Apr 20, 2024Updated 2 years ago
- ☆12Feb 21, 2021Updated 5 years ago
- ☆18Mar 10, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆51Feb 11, 2025Updated last year
- Code used for "Training Agents to Self-Report Misbehavior"☆18Feb 27, 2026Updated 6 months ago
- Reference implementation for Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model☆45Oct 1, 2025Updated 11 months ago
- The official repo for "VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search" [EMNLP25]☆39Feb 1, 2026Updated 7 months ago
- ☆16Apr 11, 2022Updated 4 years ago
- Official implementation of ICLR 2026 paper "LUMINA: Detecting Hallucinations in RAG System with Context–Knowledge Signals"☆19Jan 31, 2026Updated 7 months ago
- Repository for "Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics" …☆25Jun 16, 2024Updated 2 years ago
- ☆22Apr 27, 2026Updated 4 months ago
- ☆20Sep 16, 2025Updated 11 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Find informative examples to efficiently (human)-evaluate NLG models.☆17Apr 22, 2026Updated 4 months ago
- Official repo for "Binary Retrieval-augmented Reward Mitigates Hallucinations"☆16Nov 13, 2025Updated 9 months ago
- Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF☆25Oct 8, 2024Updated last year
- An Autonomous Curriculum Reinforcement Learning framework that steers agents to continually learn in specific environments with zero huma…☆44Jun 7, 2026Updated 2 months ago
- ☆10Jul 13, 2024Updated 2 years ago
- Code to reproduce results of our experiments using LoRe☆18Jun 10, 2026Updated 2 months ago
- rl from zero pretrain, can it be done? yes.☆296Sep 28, 2025Updated 11 months ago