The reproduct of the paper - Aligner: Achieving Efficient Alignment through Weak-to-Strong Correction
☆21May 29, 2024Updated 2 years ago
Alternatives and similar repositories for aligner-replication
Users that are interested in aligner-replication are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2024 Oral] Aligner: Efficient Alignment by Learning to Correct☆196Jan 16, 2025Updated last year
- ☆28Aug 30, 2023Updated 3 years ago
- Reimplementation of https://github.com/montemac/algebraic_value_editing in pure PyTorch for efficiency on large models☆11Jun 28, 2023Updated 3 years ago
- Official implementation of "OffsetBias: Leveraging Debiased Data for Tuning Evaluators"☆26Sep 11, 2024Updated last year
- The code and data for the paper JiuZhang3.0☆49May 26, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆13Nov 11, 2022Updated 3 years ago
- ☆16May 22, 2025Updated last year
- ☆18Jun 14, 2023Updated 3 years ago
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆14Jun 28, 2025Updated last year
- Challenge LLMs to Reason About Reasoning: A Benchmark to Unveil Cognitive Depth in LLMs☆53Jul 10, 2024Updated 2 years ago
- Direct preference optimization with f-divergences.☆17Nov 3, 2024Updated last year
- The official implementation of Self-Exploring Language Models (SELM)☆63Jun 4, 2024Updated 2 years ago
- code for "Generative News Recommendation"☆15May 31, 2024Updated 2 years ago
- ☆20Sep 16, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is the official repository of the EMNLP 2023 paper Reading Order Matters: Information Extraction from Visually-rich Documents by Tok…☆18Mar 15, 2024Updated 2 years ago
- Tools for content datamining and NLP at scale☆45Jun 20, 2024Updated 2 years ago
- GSM-Plus: Data, Code, and Evaluation for Enhancing Robust Mathematical Reasoning in Math Word Problems.☆67Jul 8, 2024Updated 2 years ago
- ☆16Apr 28, 2023Updated 3 years ago
- ☆53Apr 17, 2022Updated 4 years ago
- Improving Math reasoning through Direct Preference Optimization with Verifiable Pairs☆21Mar 20, 2025Updated last year
- The OlymMATH dataset☆24Jun 1, 2025Updated last year
- Official repository for ORPO☆479May 31, 2024Updated 2 years ago
- Implementation of the ICML 2024 paper "Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning" pr…☆117Feb 9, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆30Feb 27, 2023Updated 3 years ago
- ☆19Jul 16, 2020Updated 6 years ago
- DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents☆24Aug 4, 2025Updated last year
- Decoupled Neural Interfaces Using Synthetic Gradients - under develeopment☆11Jun 27, 2025Updated last year
- ☆29Jan 23, 2024Updated 2 years ago
- Improve Devcontainer Creation☆16May 17, 2026Updated 3 months ago
- Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation. Bingchen Zhao and Kai Han. (NeurIPS 2021)☆12Aug 20, 2023Updated 3 years ago
- This repository is about our work "A Three-Stage Self-Training Framework for Semi-Supervised Semantic Segmentation"☆20Jul 4, 2022Updated 4 years ago
- [NAACL 2025] The official implementation of paper "Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language M…☆28Mar 14, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆12Jun 22, 2023Updated 3 years ago
- The OpenAI Whisper speech-to-text model as a simple HTTP server☆14Oct 26, 2023Updated 2 years ago
- Open-Pandora: On-the-fly Control Video Generation☆35Nov 28, 2024Updated last year
- ☆86Jun 2, 2026Updated 2 months ago
- ☆14Aug 6, 2026Updated 3 weeks ago
- This is a TensorFlow implementation of DeepMind's A Distributional Perspective on Reinforcement Learning.(C51-DDPG)☆11Sep 14, 2017Updated 8 years ago
- ☆14Oct 31, 2023Updated 2 years ago