Code of "Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment" (2025).
☆14Apr 4, 2025Updated last year
Alternatives and similar repositories for regularized-bon
Users that are interested in regularized-bon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [EMNLP 2024] Introducing Filtered Direct Preference Optimization (fDPO) that enhances language model alignment with human preferences by …☆16Nov 27, 2024Updated last year
- ☆17Jun 14, 2023Updated 3 years ago
- Code for magnetic mirror descent.☆20Oct 5, 2023Updated 2 years ago
- ☆19Jun 3, 2024Updated 2 years ago
- The source code for "A Simple Graph Contrastive Learning Framework for Short Text Classification"☆13Aug 14, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Nov 20, 2025Updated 8 months ago
- Demonstrating the usage of FGYM: A Toolkit for benchmarking FPGA-accelerated Reinforcement Learning☆14Aug 12, 2021Updated 4 years ago
- docker for UTH-BERT: https://ai-health.m.u-tokyo.ac.jp/uth-bert☆14Mar 24, 2023Updated 3 years ago
- PyTorch implementation of Count-Based Exploration with Neural Density Models☆10Mar 22, 2018Updated 8 years ago
- The entire year on a single page☆12Dec 5, 2025Updated 8 months ago
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs☆17Jul 6, 2025Updated last year
- ☆15Updated this week
- Public WotLK 3.3.5a bot in C#/WPF. API surface ported from Honorbuddy, retargeted at build 12340 and custom servers. │ Botbases, navig…☆20Updated this week
- [ICML 2025] Code for "R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts"☆19Mar 10, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13May 11, 2024Updated 2 years ago
- Repository for "Propagating Knowledge Updates to LMs Through Distillation" (NeurIPS 2023).☆27Aug 25, 2024Updated last year
- Decentralized Reinforcment Learning: Global Decision-Making via Local Economic Transactions (ICML 2020)☆43Dec 8, 2022Updated 3 years ago
- MultiboxBot is a bot for multiboxing on WoW with up to 40 accounts using DLL injection, hooking and sockets.☆18Jul 8, 2026Updated last month
- ☆10Feb 12, 2026Updated 5 months ago
- Thesis project about Visual Anomaly Detection based on Self Supervised Learning. The model identifies anomalies from information acquired…☆10Apr 14, 2023Updated 3 years ago
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆15Apr 22, 2026Updated 3 months ago
- ☆13Jul 2, 2025Updated last year
- ☆44Sep 19, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- 一个用于课程小论文排版的LaTeX模板。☆10Oct 21, 2019Updated 6 years ago
- The official Python SDK for FastLabel API, the Data Platform for AI☆16Updated this week
- Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment☆62Aug 30, 2024Updated last year
- CE-GPPO: Controlling Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning☆16Jan 23, 2026Updated 6 months ago
- ☆21Sep 24, 2020Updated 5 years ago
- 更纯粹、更高压缩率的Tokenizer in Rust☆14Dec 21, 2024Updated last year
- ☆33Oct 2, 2025Updated 10 months ago
- Python library for solving Extensive-form Games and implementation of various baseline algorithms (e.g., Counterfactual Regret Minimizati…☆37Feb 12, 2026Updated 5 months ago
- ウェブサイト「サンプルで学ぶ Go 言語」のソースコード☆17Aug 17, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A repository for awesome resources in mechanistic interpretability☆16Jan 18, 2023Updated 3 years ago
- SleepCoacher recommendation engine, reference implementation for paper at http://jeffhuang.com/Final_SleepCoacher_UIST16.pdf☆26Oct 6, 2016Updated 9 years ago
- ☆21May 4, 2026Updated 3 months ago
- 해커그라운드 해커톤 2024☆12Aug 26, 2024Updated last year
- Python Vector Search tutorial generated using gpt4☆12Mar 18, 2023Updated 3 years ago
- ☆14Nov 15, 2022Updated 3 years ago
- [EMNLP2025] Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling☆16Nov 20, 2025Updated 8 months ago