DPO, but faster 🚀
☆52Dec 6, 2024Updated last year
Alternatives and similar repositories for dpo-prefix-sharing
Users that are interested in dpo-prefix-sharing are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repository contains the joint use of CPO and SimPO method for better reference-free preference learning methods.☆59Aug 13, 2024Updated last year
- Training and evaluation code for the paper "Headless Language Models: Learning without Predicting with Contrastive Weight Tying" (https:/…☆29Apr 17, 2024Updated 2 years ago
- rabitq rust implementation☆11May 14, 2026Updated 2 months ago
- torch_remat fine-grained activation checkpointing API☆15Updated this week
- ☆27Aug 28, 2025Updated 11 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated last year
- ☆24Aug 20, 2025Updated 11 months ago
- Official code of *Virgo: A Preliminary Exploration on Reproducing o1-like MLLM*☆20May 27, 2025Updated last year
- ☆15Feb 5, 2026Updated 5 months ago
- ☆37Jul 16, 2025Updated last year
- See https://github.com/cuda-mode/triton-index/ instead!☆11May 8, 2024Updated 2 years ago
- Latent Large Language Models☆19Aug 24, 2024Updated last year
- Impact of typos and common misspellings on LLM task performance.☆25Mar 22, 2024Updated 2 years ago
- Nearest Neighbor Normalization (EMNLP 2024)☆21Nov 1, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment☆16Apr 13, 2022Updated 4 years ago
- Simple repository for training small reasoning models☆51Feb 17, 2026Updated 5 months ago
- Experimental repository for research implementation of NoLoCo.☆32Mar 26, 2026Updated 4 months ago
- ☆16Feb 6, 2024Updated 2 years ago
- ☆56Apr 30, 2025Updated last year
- Official implementation of "GPT or BERT: why not both?"☆64Jul 28, 2025Updated last year
- Evaluating Reward Models in Multilingual Settings (ACL Main '25)☆43May 16, 2025Updated last year
- ☆22Jan 23, 2026Updated 6 months ago
- ☆18Apr 23, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆82Jun 8, 2026Updated last month
- ☆116Jan 21, 2025Updated last year
- Investigating the generalization behavior of LM probes trained to predict truth labels: (1) from one annotator to another, and (2) from e…☆33May 23, 2024Updated 2 years ago
- HPLT Analytics☆15Jul 22, 2026Updated last week
- Bridge Megatron-Core to Hugging Face/Reinforcement Learning☆228Jun 15, 2026Updated last month
- Contains my experiments with the `big_vision` repo to train ViTs on ImageNet-1k.☆22Jan 16, 2023Updated 3 years ago
- Official implementation of ECCV24 paper: POA☆24Aug 8, 2024Updated last year
- Accelerate LLM preference tuning via prefix sharing with a single line of code☆52Jul 4, 2025Updated last year
- Large language models (LLMs) made easy, EasyLM is a one stop solution for pre-training, finetuning, evaluating and serving LLMs in JAX/Fl…☆78Aug 17, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization☆54Jul 15, 2025Updated last year
- ☆19Nov 4, 2025Updated 8 months ago
- Landing repository for the paper "Predicting the Order of Upcoming Tokens Improves Language Modeling"☆48May 13, 2026Updated 2 months ago
- [ICLR 2026] RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling☆39Feb 25, 2026Updated 5 months ago
- [ACL 2025] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models☆39Nov 4, 2025Updated 8 months ago
- [EMNLP 2025] CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward☆68Aug 10, 2025Updated 11 months ago
- [ICLR2025] γ -MOD: Mixture-of-Depth Adaptation for Multimodal Large Language Models☆45Oct 28, 2025Updated 9 months ago