AI Fundamentals

Reward Model

A model trained to predict human preference over pairs of LLM outputs, used to provide the training signal in RLHF. Quality of the reward model bounds the quality of the RLHF-tuned policy — a limitation that DPO and Constitutional AI partly address.

Related terms

Next steps

Where this fits

Reward Model is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.