CHAPTER 03 · Alignment, Turning a Text Predictor Into a Helpful Assistant · 2 / 5
Paper 1: InstructGPT and the RLHF recipe
RLHF stands for Reinforcement Learning from Human Feedback. It is a three-step pipeline. Let us walk through each step.
flowchart TD
Base[Raw pre-trained model<br/>a good autocomplete] --> S1
subgraph S1[Step 1: Show good examples]
D[Humans write ideal answers<br/>to many prompts] --> SFT[Fine-tune the model<br/>to imitate them]
end
SFT --> S2
subgraph S2[Step 2: Learn human taste]
R[Humans rank several<br/>model answers, best to worst] --> RM[Train a reward model<br/>that scores any answer]
end
RM --> S3
subgraph S3[Step 3: Practice for a high score]
P[Model writes answers,<br/>reward model grades them] --> Opt[Model adjusts to earn<br/>higher scores]
end
Opt --> Final[Helpful, aligned assistant]
Step 1: Supervised fine-tuning (show it what good looks like)
Humans write high-quality answers to a collection of prompts. The model is then trained to imitate these examples. This step is called supervised fine-tuning, or SFT. It is like an apprentice copying a master. After this step the model is already much more helpful, but humans cannot write enough examples to cover everything.
Step 2: Train a reward model (teach the machine our taste)
This is the clever part. Instead of writing more answers, humans now just compare them. The model produces several answers to a prompt, and a person ranks them from best to worst. Ranking is much faster and more reliable for humans than writing.
These rankings are used to train a separate model called a reward model. Its only job is to look at any answer and output a score that predicts how much a human would like it. In effect, we have bottled human judgment into a piece of software that can grade an unlimited number of answers automatically.
Step 3: Reinforcement learning (practice to score well)
Now the main model practices. It writes an answer, the reward model grades it, and the model nudges itself to produce answers that earn higher scores. This is reinforcement learning, and the specific method used is called PPO. Over many rounds, the model gets better and better at pleasing the reward model, which stands in for pleasing humans.
There is one danger here. If the model only chases a high score, it might find weird tricks that fool the reward model while producing nonsense, the way a student might game a test. To prevent this, the training adds a leash called a KL penalty, which discourages the model from drifting too far from its sensible Step 1 self. Reward on one side, leash on the other, keeps it both helpful and grounded.
The headline result
The most famous finding: a small InstructGPT model with 1.3 billion parameters was preferred by humans over the original GPT-3 with 175 billion parameters, more than 100 times larger. Alignment, not raw size, was the missing ingredient. This is why every serious chat assistant since has used some form of this pipeline.