Skip to slide
Chapter 3 · Alignment, Turning a Text Predictor Into a Helpful Assistant
23 / 74

CHAPTER 03 · Alignment, Turning a Text Predictor Into a Helpful Assistant · 4 / 5

Putting the two papers together

AspectInstructGPT / RLHFDPO
Core ideaLearn human taste, then practice against itLearn directly from preferred vs rejected pairs
Reward modelSeparate, must be trainedNone, the model itself plays that role
Reinforcement learning loopYes, using PPONo
ComplexityHigh, can be unstableLow, more stable
Both rely onHuman preference dataHuman preference data

Notice the bottom row. Both methods are powered by the same fuel: humans comparing answers. The difference is only in the machinery that turns those comparisons into a better model.

← → arrow keys work too