CHAPTER 06 · Glossary: Planning and Reasoning
Glossary: Planning and Reasoning
This is your shared reference for folder 02. Every term the chapters link to is explained here from scratch, in plain language, with enough depth to actually understand it rather than just recognize it. Read it straight through as a primer, or jump in whenever a chapter sends you here.
A few foundational terms (like how a Transformer works, or what parameters are) live in the folder 01 glossary. This glossary focuses on the ideas specific to reasoning and planning.
Terms are grouped by theme so related ideas sit together.
- How models think: Prompt and prompting, In-context learning, Chain of thought, Reasoning, Reasoning trace, Self-consistency, Token, Emergent ability
- Acting in the world: Agent, Tool use, Observation and action, Hallucination, Grounding
- Checking the work: Outcome supervision vs process supervision, Verifier, Process reward model (PRM), Outcome reward model (ORM), Reward model
- Learning to reason: Reinforcement learning, Reward and reward signal, Verifiable rewards, Supervised fine-tuning (SFT), Cold-start data, Distillation
- Working at scale: Inference-time compute, Context window and long context, REPL, Recursion