TAI BUI
← Glossary
Glossary

What is RLHF?

A training pipeline: (1) collect human preferences on model outputs, (2) train a reward model on those preferences, (3) use PPO to optimize the LLM to produce higher-reward outputs

What people say

How they make AI helpful

Why it's called that