CHAPTER 02 · How Context Goes Wrong: Rot, Pollution, and Confusion · 2 / 7
Context rot: the most important failure to understand
Context rot is the phenomenon where a model's performance degrades as the context window fills up, even when the total token count is well within the technical limit.
This is a subtle and counterintuitive point, so it is worth slowing down. A model might be advertised as supporting a 1 million token context window. You might reasonably assume that as long as you stay under 1 million tokens, the model performs at full strength. The articles say this is false. The "effective context window," the region where the model actually reasons well, is often far smaller, frequently under 256,000 tokens for current models. Past that, quality quietly drops even though no error is thrown.
The article points to outside evidence for this. Chroma published a study on context rot, and Anthropic has explained that a growing context depletes the model's "attention budget." A helpful way to picture it: attention is a limited resource that gets spread across everything in the window. The more you cram in, the thinner that attention is spread, and the easier it becomes for the model to miss the part that actually mattered.
The practical consequence is a rule the article calls the "Pre-Rot Threshold." Do not wait for the API to throw an error at the hard limit. Instead, monitor your token count and start cleaning up the context well before the rot zone, for example by triggering compaction or summarization at a chosen threshold. The cleanup techniques are covered in Chapter 3.