TAI BUI
← Glossary
Glossary

What is DPO?

A training method that skips the reward model entirely — it directly optimizes the language model to prefer the better response in pairs of human preferences

What people say

A simpler RLHF

Why it's called that