← Glossary
Glossary
What is DPO?
A training method that skips the reward model entirely — it directly optimizes the language model to prefer the better response in pairs of human preferences
What people say
A simpler RLHF
A training method that skips the reward model entirely — it directly optimizes the language model to prefer the better response in pairs of human preferences
A simpler RLHF