☆16Jul 23, 2024Updated 2 years ago
Alternatives and similar repositories for refdpo
Users that are interested in refdpo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation for the paper "Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning"☆11Jan 10, 2025Updated last year
- the datasets of our paper☆11Feb 26, 2024Updated 2 years ago
- The official repository of "Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint"☆39Jan 12, 2024Updated 2 years ago
- SIFT: Grounding LLM Reasoning in Contexts via Stickers☆58Mar 6, 2025Updated last year
- This repository includes code and materials for the paper "Efficient PRM Training Data Synthesis via Formal Verification" (ACL 2026 Findi…☆19Apr 7, 2026Updated 4 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This is the oficial repository for "Safer-Instruct: Aligning Language Models with Automated Preference Data"☆17Feb 22, 2024Updated 2 years ago
- [NeurIPS 2024] Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study☆63Nov 24, 2024Updated last year
- ☆25Dec 13, 2024Updated last year
- A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data☆36May 1, 2026Updated 3 months ago
- ☆12Apr 25, 2025Updated last year
- ☆37Sep 16, 2024Updated last year
- Codebase for Instruction Following without Instruction Tuning☆36Sep 24, 2024Updated last year
- The Code and Script of "David's Slingshot: A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis"☆34Jun 13, 2025Updated last year
- Co-LLM: Learning to Decode Collaboratively with Multiple Language Models☆129May 7, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Universal Neurons in GPT2 Language Models☆30May 28, 2024Updated 2 years ago
- MATS: A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback☆18Dec 27, 2025Updated 7 months ago
- Implementation for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"☆397Jan 19, 2025Updated last year
- Official code and data repository of MathChat: MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Inte…☆22Jun 3, 2024Updated 2 years ago
- [ACL 2026 Oral] Official implementation of LaMI: Augmenting Large Language Models via Late Multi-Image Fusion☆19Jul 4, 2026Updated last month
- ☆15Apr 11, 2024Updated 2 years ago
- Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment☆62Aug 30, 2024Updated last year
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆33Jan 23, 2025Updated last year
- ACL24☆11Jun 7, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- ☆15Feb 10, 2026Updated 6 months ago
- The code and data for the paper JiuZhang3.0☆49May 26, 2024Updated 2 years ago
- ☆34Sep 14, 2024Updated last year
- The code for "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by VIdeo SpatioTemporal Augmentation" [CVPR2025]☆20Feb 27, 2025Updated last year
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 9 months ago
- ☆23Jul 5, 2024Updated 2 years ago
- ☆16Jul 10, 2023Updated 3 years ago
- ☆12Aug 6, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- LMM for VQA, tcsvt version☆10Jul 19, 2024Updated 2 years ago
- AdaICL: Which Examples to Annotate of In-Context Learning? Towards Effective and Efficient Selection☆19Oct 30, 2023Updated 2 years ago
- Official repository for paper: O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning☆100Feb 21, 2025Updated last year
- Official implementation of Bootstrapping Language Models via DPO Implicit Rewards☆48Apr 15, 2025Updated last year
- A web app for both Text-based and Visual Question Answering.☆13Nov 13, 2023Updated 2 years ago
- Dateset Reset Policy Optimization☆30Apr 12, 2024Updated 2 years ago
- FocusLLM: Scaling LLM’s Context by Parallel Decoding☆45Dec 8, 2024Updated last year