The reproduct of the paper - Aligner: Achieving Efficient Alignment through Weak-to-Strong Correction
☆21May 29, 2024Updated 2 years ago
Alternatives and similar repositories for aligner-replication
Users that are interested in aligner-replication are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 33B Chinese LLM, DPO QLORA, 100K context, AirLLM 70B inference with single 4GB GPU☆14May 5, 2024Updated 2 years ago
- [NeurIPS 2024 Oral] Aligner: Efficient Alignment by Learning to Correct☆196Jan 16, 2025Updated last year
- [EMNLP2023]: MIRACLE: Towards Personalized Dialogue Generation with Latent-Space Multiple Personal Attribute Control☆12Nov 11, 2023Updated 2 years ago
- ☆28Aug 30, 2023Updated 3 years ago
- ☆13Nov 11, 2022Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Self-hosted LLM chatbot arena, with yourself as the only judge☆41Feb 6, 2024Updated 2 years ago
- ☆16May 22, 2025Updated last year
- ☆17Oct 18, 2022Updated 3 years ago
- Complexity Based Prompting for Multi-Step Reasoning☆17Mar 10, 2023Updated 3 years ago
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆14Jun 28, 2025Updated last year
- Challenge LLMs to Reason About Reasoning: A Benchmark to Unveil Cognitive Depth in LLMs☆53Jul 10, 2024Updated 2 years ago
- Direct preference optimization with f-divergences.☆17Nov 3, 2024Updated last year
- The official implementation of Self-Exploring Language Models (SELM)☆63Jun 4, 2024Updated 2 years ago
- 4 bits quantization of LLaMa using GPTQ☆12Jun 2, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆16Oct 5, 2022Updated 3 years ago
- ICML 2024 - Official Repository for EXO: Towards Efficient Exact Optimization of Language Model Alignment☆55Jun 16, 2024Updated 2 years ago
- ☆20Sep 16, 2025Updated last year
- Beginner-friendly serverless LLM deployment with Replicate & fly.io☆13Sep 3, 2023Updated 3 years ago
- GSM-Plus: Data, Code, and Evaluation for Enhancing Robust Mathematical Reasoning in Math Word Problems.☆67Jul 8, 2024Updated 2 years ago
- Universal LLM Telegram chatbot in Python☆17Aug 16, 2024Updated 2 years ago
- ☆53Apr 17, 2022Updated 4 years ago
- ☆118Jan 21, 2025Updated last year
- Improving Math reasoning through Direct Preference Optimization with Verifiable Pairs☆21Mar 20, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆17Sep 30, 2023Updated 2 years ago
- The OlymMATH dataset☆25Jun 1, 2025Updated last year
- Official repository for ORPO☆480May 31, 2024Updated 2 years ago
- Implementation of the ICML 2024 paper "Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning" pr…☆117Feb 9, 2024Updated 2 years ago
- DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents☆24Aug 4, 2025Updated last year
- Xwin-LM: Powerful, Stable, and Reproducible LLM Alignment☆1,037May 31, 2024Updated 2 years ago
- Improve Devcontainer Creation☆17Sep 3, 2026Updated 2 weeks ago
- Official code for "MAmmoTH2: Scaling Instructions from the Web" [NeurIPS 2024]☆146Oct 27, 2024Updated last year
- Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation. Bingchen Zhao and Kai Han. (NeurIPS 2021)☆12Aug 20, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- This repository is about our work "A Three-Stage Self-Training Framework for Semi-Supervised Semantic Segmentation"☆20Jul 4, 2022Updated 4 years ago
- [NAACL 2025] The official implementation of paper "Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language M…☆28Mar 14, 2024Updated 2 years ago
- ExploitGen is a template-augmented Exploit Code generation method based on CodeBERT, which accepted in JSS.☆11Feb 9, 2024Updated 2 years ago
- team Doggeee's solution to Ego4D LTA challenge@CVPRW23'☆14Nov 4, 2023Updated 2 years ago
- ☆22Jan 19, 2022Updated 4 years ago
- Open-Pandora: On-the-fly Control Video Generation☆35Nov 28, 2024Updated last year
- ☆14Oct 31, 2023Updated 2 years ago