papers related to Direct Preference Optimization(DPO)
☆20Jul 16, 2024Updated 2 years ago
Alternatives and similar repositories for awesome-DPO
Users that are interested in awesome-DPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [WWW 2023] The official code for the paper "Two-Stage Constrained Actor-Critic for Short Video Recommendation"☆15Jul 21, 2023Updated 3 years ago
- [ECCV Workshop W-CODA 2024] Official code for ReGentS☆18Oct 22, 2024Updated last year
- ☆16Oct 3, 2024Updated last year
- [ICML 2023] Meta-SAGE: Scale Meta-Learning Scheduled Adaptation with Guided Exploration for Mitigating Scale Shift on Combinatorial Optim…☆11Dec 19, 2023Updated 2 years ago
- ☆18Nov 9, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- a tiny project to test the effectiveness of video QA through RAG techniques and multimodal LLMs☆15Jun 2, 2024Updated 2 years ago
- An open-source server implementation for inference Qwen2-VL series model using fastapi.☆10Nov 20, 2024Updated last year
- Interpreting Learned Search and Planning: Reverse-engineering recurrent convolutional networks (DRC) that play Sokoban☆22Jun 29, 2025Updated last year
- From-Classification-to-Clinical☆13Apr 26, 2024Updated 2 years ago
- Record experiment data easily☆14Aug 13, 2022Updated 4 years ago
- particle filter based object tracking☆17Mar 9, 2020Updated 6 years ago
- Data for evaluating GPT-4V☆11Oct 26, 2023Updated 2 years ago
- Stores here are the source codes for the official implementation of "Generating Traffic Scenarios via In-Context Learning to Learn Better…☆25May 1, 2025Updated last year
- ☆27Apr 14, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A list of papers regarding generalization in (deep) reinforcement learning☆11Aug 13, 2023Updated 3 years ago
- "AI Commit Message Tool uses AI to automatically generate concise and professional Git commit messages, which you can then edit and confi…☆14Jul 14, 2025Updated last year
- Python and OpenCV program to estimate Fundamental and Essential matrix between successive frames to estimate the rotation and the transla…☆22May 21, 2019Updated 7 years ago
- Codebase for Paper Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs☆24Apr 24, 2025Updated last year
- Get data from Tobii Eye Tracker 4C/5 with Python☆18Sep 19, 2022Updated 4 years ago
- [NeurIPS‘2021] "MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge", Geng Yuan, Xiaolong Ma, Yanzhi Wang et al…☆18Mar 16, 2022Updated 4 years ago
- NLP☆14Oct 17, 2022Updated 3 years ago
- ☆19Oct 8, 2024Updated last year
- The official implementation of Preference Data Reward-Augmentation.☆18May 1, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACMMM 2025] Official Code of DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Model…☆27Sep 23, 2025Updated 11 months ago
- ☆16May 22, 2025Updated last year
- MagicRF M100/QM100☆29Oct 29, 2021Updated 4 years ago
- ☆20Aug 14, 2025Updated last year
- ☆19Nov 11, 2024Updated last year
- [CVPR25 Highlight] A ChatGPT-Prompted Visual hallucination Evaluation Dataset, featuring over 100,000 data samples and four advanced eval…☆32Apr 16, 2025Updated last year
- SDD 规格驱动开发 + Harness 多 Agent 编排;支持在 OpenCode、Claude Code、Codex CLI 间切换执行引擎,完成从规格到落地的自动化开发任务。☆15Apr 8, 2026Updated 5 months ago
- FairGAN: GANs-based Fairness-aware Learning for Recommendations with Implicit Feedback☆15Oct 8, 2022Updated 3 years ago
- ☆35Aug 19, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LightRFT (Light Reinforcement Fine-Tuning) is an advanced reinforcement learning fine-tuning framework designed for Large Language Models…☆19Jan 12, 2026Updated 8 months ago
- [ACL 2025 Findings] Official implementation of the paper "Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning".☆23Feb 26, 2025Updated last year
- Code for "ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch", where dataset…☆17Sep 8, 2025Updated last year
- Implementation of paper "Do Wide and Deep Networks Learn the Same Things?"☆16Mar 15, 2022Updated 4 years ago
- ☆17Jun 3, 2025Updated last year
- ICS_2020_PJ☆11Dec 25, 2020Updated 5 years ago
- This repository hosts the DataAssistant, a robust Python class designed to integrate seamlessly with OpenAI's API. It facilitates the cre…☆13Jul 2, 2024Updated 2 years ago