AI Fundamentals

DPO

Direct Preference Optimization — an RLHF-like fine-tuning approach that skips training an explicit reward model, instead directly optimizing the policy against pairwise preferences. Simpler and often more stable than PPO-based RLHF; widely adopted in open-weights fine-tuning.

Related terms

Next steps

Where this fits

DPO is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.